Pangram verdict · v3.3
We believe this text is mainly AI, with some human-written content.
AI likelihood · overall
AIArticle text · 655 words · 3 segments analyzed
PublicationsHow open-weight demand can support the useful life of NVIDIA GPU familiesOrnn Data7 September 2026AbstractGPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete.
In this paper, we provide a counter to this thesis by examining the effect of open-weight demand on the economic usefulness of older GPU families. Closed-model access runs through subscription allowances that the provider is able to reset, so the posted token rate card represents the marginal price of additional usage. Across eleven open-weight and eight closed models on the Artificial Analysis Intelligence Index, the cheapest qualifying open-weight model, standardized by intelligence, completes a task at roughly one fifth of the cost of a comparable closed model. Self-hosting on rented hardware lowers this to $0.12 to $0.35 per million output tokens at full utilization and reverses the hardware ranking: on gpt-oss-120b, a sparse model with 5.1 billion active parameters, the A100 produces output more cheaply than the H100 at spot and at the three- and five-year term prices. Ornn’s rental data show the market reflecting this utility. The five-year A100 rental price maintains 80 percent of its one-month term price (vs 44 to 60 percent for the Hopper and Blackwell families) for a contract ending when the Ampere family is more than eleven years old. We show that today’s compute-intensive workloads—long-running agents, batch evaluation, and reinforcement learning—tolerate latency and are hardware agnostic, which incentivizes price-elastic demand to route to any cost-efficient hardware. See NVIDIA’s acquisition of Hugging Face on 3 September 2026 (NVIDIA, 2026c). These findings challenge forecasts that newer hardware eliminates the earning capacity of older GPUs. Instead, they suggest that older NVIDIA generations retain a multi-year earning life so long as they serve suitable workloads competitively and operators remain free to deploy those workloads on them.Key findingsHosted open-weight models were cheaper at several sampled common score thresholds on the Artificial Analysis Intelligence Index, not at every threshold. Closed models remain the cheaper qualifying option at some scores, and the closed frontier exceeds the open sample at the top of the range.Self-hosting on rented hardware produced compute-only costs of $0.12 to $0.35 per million output tokens at full utilization in the printed sample.On gpt-oss-120b, the A100 produced output more cheaply than the H100 at spot and at the three- and five-year term prices. The A100/H100 sparse result combines different third-party serving setups. The dense A100 row is estimated.The five-year A100 term price retains 80.2% of the one-month price, versus 43.7% to 59.8% for Hopper and 53.8% for Blackwell. Forward marks are analyst-produced indicators, not executable quotes.A100 occupancy rose from 74% to 90% as listed capacity rose 13% and the spot index rose 20%. The paper does not establish that open-weight demand caused A100 occupancy or rental-price behavior.Selected figures and tables$0.00$0.50$1.000.750.29A100 SXM40.680.64H100 SXM0.810.98H2000.390.47B200Dense (Llama-2-70B)Sparse (gpt-oss-120b)Compute-only cost per million output tokens at full use and the base caseGPUDense full useDense base caseSparse full useSparse base caseA100 SXM4$0.28$0.75$0.12$0.29H100 SXM$0.26$0.68$0.27$0.64H200$0.32$0.81$0.35$0.98B200$0.16$0.39$0.15$0.47Source: Ornn Data calculation from published throughput and 1 September 2026 spot rents (paper Table 7). Full use takes Offline throughput with no headroom. The base case takes Server throughput where available and otherwise rescales the GPUStack baseline; the latter does not establish a Server service level. Dense A100 is estimated. The A100/H100 sparse rows are third-party measurements, not MLPerf results. Costs are compute-only USD per million output tokens.A100 occupancy, listed-capacity change, and spot-price changeFamilyOccupancy, 1 Mar 2026Occupancy, 1 Sep 2026Occupancy, Mar–Sep meanListed capacity, Mar–SepSpot, Mar–SepA100 SXM474%90%80%+13%+20%Source: Ornn occupancy and listed-capacity series and the settled spot index (paper Table 8), retrieved 2 September 2026. Occupancy is rented capacity divided by listed capacity across tracked on-demand providers in the global region.
Listed capacity measures tracked on-demand supply, not the installed base.406080100806044541M6M1Y3Y5YTerm price as % of the 1-month markA100 SXM4H100 SXMH200B200B300Forward term-price retention by GPU familyFamily6M / 1M1Y / 1M3Y / 1M5Y / 1MFwd 37–60 / 1MAge at 5Y end (yrs)A100 SXM495.0%93.1%84.2%80.2%74%11.3H100 SXM95.0%86.2%68.2%59.8%47%9.4H20092.4%80.2%51.0%43.7%33%7.8B20099.1%97.9%69.1%53.8%31%7.5B30093.8%81.3%60.9%53.8%43%6.5Source: Ornn reported term marks (paper Table 9).