Pangram verdict · v3.3
We believe that this entire text is AI.
AI likelihood · overall
AIArticle text · 1,511 words · 1 segments analyzed
01 / WHAT THE NUMBER MEANS “A million a second” can mean four things. CONTEXT CAPACITYHow much fitsThe text a model can hold in one request, including its reply. A size, not a speed. INPUT PROCESSINGHow fast it readsPrompt tokens can be processed largely in parallel. Huge inputs still take time, and the input rate is different from the output rate. AGGREGATE THROUGHPUTHow much a system serves10,000 streams × 100 tokens/sec = 1,000,000 tokens/sec in total. Toy arithmetic: it says nothing about when any one stream finishes. SINGLE-AGENT OUTPUT · OUR DIALHow fast one agent writesOne agent writing a million tokens a second. In standard generation, each token depends on the ones before it. Anthropic’s Claude Opus 5.5 overview, for example, lists a 1M-token context window, a much smaller output cap per request and “moderate” comparative latency. Memory size is not writing speed, and it does not mean a million-token reply fits in one request. Engineering helps, within limits. Services reach big totals by batching many requests, as the 2023 vLLM paper describes, and speculative decoding can accept several drafted tokens in one pass. Neither removes a genuine chain in which step two needs the result of step one. A token is a chunk of text, often part of a word, so token counts are not word counts. TRY IT / THE IMAGINED DIAL Turn the dial. Pick a workload, then slide the imagined speed from 100 to 1,000,000 tokens per second. Imagined speed for one agent1,000,000 tokens/sec OUTPUT TOKENS1,000,0001,000 story endings × 1,000 tokens each WRITING ONLY1 secondat 1,000,000 tokens/sec At 1,000,000 tokens/sec, 1,000 story endings (1,000,000 output tokens) take about 1 second to write. All four budgets at both ends of the dial Writing-only time at illustrative speeds, not measurements. WorkloadOutput tokens100 tokens/sec1,000,000 tokens/sec 1,000 story endings × 1,0001,000,0002 h 46 min 40 s1 second 40 app drafts × 20,000800,0002 h 13 min 20 s0.8 seconds 10,000 critiqued candidates × 3003,000,0008 h 20 min3 seconds 10,000 rehearsals × 1,00010,000,00027 h 46 min 40 s10 seconds Imagined comparisons, not measurements of any model. Times count generated output only: no reading, ranking, tool calls, permissions, tests, deployment, people or experiments. A budget is a writing allowance, not a finished product. NoticeParallel streams can accelerate these independent jobs too. But each draft still needs time on its own stream: matching aggregate throughput does not match completion time. One fast agent matters most when each step waits on the last. 02 / FOUR THOUGHT EXPERIMENTS Cheap drafts move the hard part. Each budget counts written output only. None is a finished product, a verified discovery or a real result. A map of possible endings 1,000 × 1,000 = 1,000,000 tokens · 1 second of writing Ask for an ending to your story and get a thousand: hopeful, dark, strange. Nobody reads a thousand, so the useful version maps them for you to explore, then blends the two you like. Still scarce: taste. Only you know which ending is yours. Software for a neighborhood tool library 40 × 20,000 = 800,000 tokens · 0.8 seconds of writing Describe a lending app for shared tools and forty drafts exist before you finish the sentence. Screens could adapt, from borrowing a ladder to a repair-day sign-up. Tests still run on their own clock, and if none checks the seven-day due date, all forty can pass while lending ladders for seventy. Still scarce: a clear spec, and tests of what “working” means. Ten real experiments from ten thousand ideas 10,000 × 300 = 3,000,000 tokens · 3 seconds of writing Hunting for a better catalyst? An agent could propose and critique ten thousand candidates. Then everything waits at the bench: reactions may run for hours, cultures for days, field trials for a season. Ideas from one model can also share one blind spot. Still scarce: physical evidence, and choosing which ten experiments earn lab time. Rehearsing a hard conversation 10,000 × 1,000 = 10,000,000 tokens · 10 seconds of writing Before talking to your landlord about the lease, an agent could play the conversation ten thousand ways: stubborn landlord, generous landlord, you when tired. Use it like a flight simulator, for practice and blind spots. Ten thousand rehearsals of the wrong person are a confident mistake, not a prophecy. Still scarce: fidelity to the real person, and your own practice. 03 / THE SERIAL BOTTLENECK 10,000× faster writing is not 10,000× faster work. Take the forty app drafts: 800,000 output tokens. Compare an illustrative 100 tokens per second with the imagined million, then add one fixed check after writing that speed does not touch, such as a test run. Fixed check after writing60 seconds Illustrative 100 tokens/sec2 h 14 min 20 s Writing 99.3% · Checking 0.7% Imagined 1,000,000 tokens/sec60.8 seconds Writing 1.3% · Checking 98.7% WRITING SPEEDUP10,000×1,000,000 ÷ 100 tokens/sec END-TO-END SPEEDUP132.57×8,060 ÷ 60.8 seconds With 60 seconds of checking, the batch takes 2 h 14 min 20 s at 100 tokens/sec and 60.8 seconds at 1,000,000 tokens/sec. That is 132.57× faster end to end, and checking is 98.7% of the faster total. The same 800,000 tokens with four check times Total time = writing + one fixed check, counted once. Fixed check100 tokens/sec1,000,000 tokens/secEnd to end None2 h 13 min 20 s0.8 seconds10,000× 60 seconds2 h 14 min 20 s60.8 seconds132.57× 15 min2 h 28 min 20 s15 min 1 s9.88× 24 h26 h 13 min 20 s24 h 1 s1.09× Toy model: total time = output tokens ÷ speed + one fixed check, counted once for the whole batch (not per draft) after writing ends. It does not simulate parallel tests, cost, energy or draft quality. With a one-minute check, the job drops from 8,060 to 60.8 seconds: about 132.57 times faster, not 10,000. At zero the full 10,000× returns; at a day the gain nearly vanishes. Whatever you do not speed up becomes nearly all the remaining time, the logic of Amdahl’s law. Real requests have more such steps: OpenAI’s latency guide notes that very large prompts, tool calls and network trips add delays of their own. 04 / WHAT GETS PRECIOUS When generation gets cheap, judgment gets precious. Many drafts, one narrow gate Conceptual drawing. A wide grid of draft cards on the left flows toward a narrow gate. One card passes through and is marked as chosen. It illustrates the argument, not data. Conceptual drawing, not data: generation widens the options; judgment decides what passes. Every experiment above ends in the same place. What stays scarce is judgment, wearing different hats: TasteWhich option is yours. Specs and testsWhat “working” means. PracticeSpeed cannot learn it for you. Physical evidenceThe world answers at its own pace. FidelityA simulation is only as good as its model. Cost and energyIs this worth running at all? Fast does not mean cheap: every token runs on hardware someone pays to power. More options do not guarantee a good one either. Drafts from one model and one set of assumptions can be correlated, or all wrong together. Volume measures output, not understanding. 05 / USE IT THIS WEEK Before you ask for more, decide how you will choose. No imaginary dial required. Next time you hand work to an AI agent: Write down what “done” means.One testable sentence, like “ladders are due back in seven days,” not “handles loans.” Set criteria before reading options.Two or three, so the most fluent draft does not win by default. Find the slow step.Name the check that speed will not touch. Shorten or automate it where possible, without skipping the validation each result needs. Ask for disagreement, not volume.Request options built on different assumptions, then ask what would make all of them wrong. If your slow step is deciding what to measure or which bets to make, that is strategy more than tooling. Private consulting helps clarify the larger system and where your next move matters ↗ Sources, dates and boundaries. This Field Note adapts ideas from echohive’s film One Million Tokens a Second; it is an edited companion, not a transcript. The speed is imagined, and no source below measures, claims or predicts it. They support only the distinctions between capacity, reading, serving and writing. Open the four sources Anthropic, Claude Opus 5.5 overview. Official documentation. Lists a 1M-token context window, a separate maximum output per request and “moderate” comparative latency. It does not describe a million output tokens per second. Anthropic, Context windows. Official documentation. Describes the context window as working memory that includes the generated response: a capacity, not a speed. OpenAI, Latency optimization. Official API guide. Output generation commonly dominates latency; very large prompts still matter, and tool or network calls add their own delays. Agrawal et al., Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. arXiv, 2024. A historical foundation, not a current benchmark: parallel prompt processing (prefill), token-by-token decoding and batching. The 2023 vLLM paper linked in section 01 is also a historical foundation. All workloads, speeds, budgets and the bottleneck model are illustrative arithmetic, not benchmarks, forecasts or product specifications. Sources checked October 2, 2026.