Skip to content
HN On Hacker News ↗

Tokens to GPUs Calculator | Cedana

▲ 21 points • 9 comments • by kmavm • 3d ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

66 %

AI likelihood · overall

Mixed
26% human-written 74% AI-generated
SEGMENTS · HUMAN 1 of 3
SEGMENTS · AI 2 of 3
WORD COUNT 652
PEAK AI % 84% · §1
Analyzed
Oct 7
backend: pangram/v3.3
Segments scanned
3 windows
avg 217 words each
Distribution
26 / 74%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 652 words · 3 segments analyzed

Human AI-generated
§1 AI · 84%

» GPU count367H100 GPUs / base211Low853High468x H100 nodes1 trillion tokens in one month (30 days) on Llama 3.3 70B needs approximately 367 H100 GPUs. Confidence is moderate.1101001k10kSingle GPUsNodes with 8 GPUsBase estimatePlausible rangeLog axis. Select a row to see its calculation.One model copy needs 3 V100 GPUs (77 GB against 32 GB). Treat the V100 result as theoretical.» CalculationHow the calculator gets the H100 numberEach step changes one quantity. Read from left to right.01Start1 trilliontokens in one month (30 days)02÷ 2,592,000 seconds385,802tokens per second, average03× 0.475 for the token mix183,256output-equivalent tokens per second04÷ 50% utilization366,512tokens per second of installed capacity05÷ 1,000 tok/s for each H100367H100 GPUsgpus = tokens ÷ seconds × [output share + input share × input cost] ÷ utilization ÷ tok/s for each gpu» SensitivityWhich assumption moves the H100 result mostEach bar shows the GPU count when one assumption moves from its best case to its worst case. The other assumptions stay at the base value. The longest bar is the assumption to measure first.H100 speed for this modelmeasurement uncertainty244 at 1,500 tok/s814 at 450 tok/sAverage utilizationyour operating choice229 at 80%611 at 30%Cost of an input tokenmeasurement uncertainty251 at 0.1 of output482 at 0.5 of outputMeasurement uncertaintyYour operating choiceVertical line: base, 367 H100 GPUs before rounding» Per-GPU detailGPUtok/s per GPUTokens per GPU per dayGPUs per copyLowBaseHigh8-GPU nodesA10045040.93M14368151,966102V10014413.10M31,3382,5477,122319H1001,00090.95M121136785346B2004,000363.79M1479225212» Model presets and evidenceReference speed is tokens per second for one GPU, as low, base, and high. "Measured" values come from public benchmarks across loose and strict latency targets.

§2 Human · 15%

"Estimated" values have no direct benchmark.ModelTotal / activeWeightsReference speedB200 ÷ H100BasisSourceLlama 3.1 8B8B / 8B8 GBH100: 3,000, 6,000, 12,5002.3, 4, 6Estimated. One measured point for the high value.morph »Llama 3.3 70B70B / 70B70 GBH100: 450, 1,000, 1,5002.3, 4, 6Measured on H100 and B200.inferencex »cerebrium »gpt-oss 120B117B / 5.1B65 GBH100: 740, 1,400, 2,6006, 9, 12Measured on H100 and B200. Weight size is not verified.inferencex »DeepSeek V4 Flash284B / 13B160 GBB200: 1,500, 4,000, 9,0005, 9, 14Estimated. No benchmark found. Scaled from V4 Pro and gpt-oss by active parameters.datacamp »MiniMax M3 428B428B / n/a428 GBH100: 157, 240, 5373.5, 4.5, 5.3Measured on H100 and B200. Weight size assumes 8-bit.inferencex »DeepSeek R1 / V3671B / 37B671 GBH100: 23, 75, 26612, 16, 21Measured on H100 and B200. Weight size assumes 8-bit.inferencex »GLM-5.3753B / 40B755 GBB200: 1,400, 3,000, 7,0004, 8, 16Measured on B200 only. Sources disagree: GLM-5 at FP4 gives 1,417 to 2,935, GLM-5.3 gives 4,329 to 10,243. H100 ratio is estimated.inferencex 5.3 »inferencex 5 fp4 »size »DeepSeek V4 Pro1600B / 49B865 GBB200: 468, 1,000, 2,8006, 14, 21Measured on B200 only.

§3 AI · 83%

H100 ratio is estimated from DeepSeek R1.inferencex »size »» Other assumptionsAssumptionValueBasisSourceA100 ÷ H1000.35, 0.45, 0.60H100 measured at 1.8x to 2.9x the A100 on 70B modelshyperstack »perplexity »V100 ÷ H1000.10, 0.18, 0.25Specifications only: 125 against 312 TFLOPS, 900 against 2,039 GB/s. No LLM serving benchmark.spheron »Input cost0.1, 0.3, 0.5Derived from 128:128 and 2024:128 runs on 4x H100 (result 0.33)e2e networks »GPU memory32, 80, 80, 180 GBV100, A100, H100, B200. The fit check adds 10% to the weight size.inferencex »Values read on 2 October 2026.» Limits of this modelThe result is a planning estimate. Expect an error of 2x in each direction for measured presets, and more for estimated presets. Measure your model on your hardware before you buy.Speed changes with the latency target. The low and high values of each preset come from strict and loose latency targets.For large mixture-of-experts models, the B200 lead over the H100 is 4x to 21x. FP4 support and memory size cause this. The ratio is specific to each model.A100 and V100 results for the large models are theoretical. One model copy needs more GPUs than one node contains.The V100 value has no measured LLM serving data. It comes from hardware specifications.The range combines independent errors as a root sum of squares in log space. It is not a statistical confidence interval.The calculator does not include prompt caching, speculative decoding, reasoning-token overhead, failures, or regional duplication.