Skip to content
HN On Hacker News ↗

tokens too cheap to meter

▲ 354 points • 227 comments • by teoruiz • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,698
PEAK AI % 0% · §1
Analyzed
Sep 23
backend: pangram/v3.3
Segments scanned
1 windows
avg 1698 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,698 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

The price of using machine learning intelligence is decreasing by several orders of magnitude a year and shows no signs of slowing. We are likely to see LLMs integrated into every part of computing as infrastructure, not just as a product, in the next year or two. We are likely to see LLMs running locally at current frontier-quality on commodity hardware in the next 3-6 years. Starting very soon, we are likely to see quality and access become the limiting factor to AI 1 use, not sheer number of tokens. Is this really happening? Extraordinary claims require extraordinary evidence, so I collected a whole bunch of evidence. AI can be either proprietary (such as GPT-6 Astra) or open weight (such as GLM-5.3-flash). Open weight models can be either hosted (e.g. by Z.ai) or local. Generally, models intended to be run locally will be much smaller, such as Muse Glimmer or Qwen3 Coder. Improvements in one don't always affect improvements in the others. Improvements that affect all AI GPUs GPUs are getting exponentially more efficient with every generation. In the graph below (source), the X-axis is time and the Y-axis is power efficiency of the GPU itself. Larger Y-axis numbers mean more efficient. This is a logarithmic graph, which is to say that a straight line on the graph represents an exponential increase in efficiency. In this particular case, the logarithm is 1.3, which means efficiency doubles about once every two years. This is an increase in efficiency that we haven't seen since Moore's Law in the 1960s. Models The cost to complete a given task with a model is going down sharply over time. Models are usually priced per-token. A "token" is a fragment of a word; it takes about 1.5 tokens to represent a word. For every token a model reads, and for every token it outputs, the "model provider" (e.g. Anthropic or OpenAI) charges you some fixed amount of money. The cost per token of models is not consistently going down, at least not for the smartest ("frontier") models. But the cost per task is. Smaller models may cost less per token, but use more tokens overall than a larger model for the same task, because they have to think more or correct their first drafts. This section is about the cost to complete the task from beginning to end. The chart below (source) shows the "pareto frontier" of cost/task at present. A pareto frontier shows the best tradeoff you can get, not just the best in a single category. Here, our tradeoffs are: Y-axis: the "quality" of the model (as measured by a suite of benchmarks) X-axis: the cost to complete those benchmarks Cost is on a logarithmic scale. Larger Y-axis and smaller X-axis numbers are better. This is showing us a wide range of models on the pareto frontier as of 2026. Towards the top-right we have Claude Fable-5.1 (expensive and intelligent); towards the middle-left we have GPT-5.6 Luna (cheap and less intelligent). Models below the dotted line are basically not worth considering.2 Now, look at this chart showing the frontier at the start, middle, and end of 2025: The chart shows models are getting smarter and cheaper on a per-task basis over 2025. If you draw a straight horizontal line at basically any task on the Y-axis, the cost to do it at the end of 2025 was cheaper than at the start; and if you draw a straight vertical line at basically any point on the X-axis, models can do more for the same cost. Now, compare that 2025 chart to the 2026 chart. The Y-axis (intelligence) is about the same, with less of a fall-off towards the cheap end. The X-axis (cost) has gotten two orders of magnitude cheaper. Inference Engines An "inference engine" is a software package that takes a trained model and an input text and actually runs it on a GPU. Inference engines are currently immature and improving rapidly. Currently we're seeing 10%-50% improvements year-over-year, depending on which engine you look at. There are two benchmarks that are often compared for inference engines: "offline" (run a bunch of tokens through in one big batch) and "serving" (you have people sending your server inputs at unpredictable times, and you want to send a response back as quickly as possible). Serving is getting efficient much more rapidly than offline inference. All numbers below are for serving workloads, not offline. vLLM vLLM is an open-source inference engine and it's getting more efficient over time. In the graph below (source), the Y-axis is Joules/token, the X-axis is batch size (roughly: "how many inputs are processed in parallel?"), and the blue/red lines are different software versions. Smaller Y-axis numbers mean more efficient. vLLM 0.11.1 was released in December 2025, a bit more than a year after vLLM 0.5.4 in September 2024. In other words, this is about a 40% increase in efficiency in 15 months. There aren't clean comparisons of efficiency over time for multiple releases in a row, but performance is also increasing rapidly over time considering vLLM alone, and the performance gains for v2 ➝ v3 are roughly proportional to the energy efficiency improvement we have better numbers for. NVIDIA This isn't isolated to a single software package. NVIDIA is showing up to 50% efficiency improvements on their MLPerf stack from 2.0 to 2.1: Intel This isn't isolated to old benchmarks. Intel recently showed a 2.4x throughput increase solely by improving MLPerf between 6.0 and 6.1. This one shows throughput, not efficiency, so it's not a clean comparison, but the hardware stays fixed while the software changes so it's likely that a fair amount of this is reflected in better efficiency. Improvements that affect hosted AI Mixture-of-Experts Models are using architectures that are fundamentally more efficient than early ways we knew how to build an LLM. Early LLMs were based around "dense" models. This means that every part of the model is "activated" (runs a matrix multiplication) on every input. Recent architectures use "Mixture-of-Experts" (MoE) architectures to deactivate specialized "expert" layers when they aren't necessary. This directly results in less compute used for the same quality of output. In the graph below, a model can be 7x smaller (6B ➝ 0.8B parameters) while achieving the same performance on benchmarks (source): This means we're going to see the cost and memory usage of models go down over time, relative to the quality of the model. Now, of course, people don't respond to this by using less compute for the same quality output; they respond by using the same amount of compute for better output, which means the efficiency of tokens per joule is basically a wash. However, the efficiency of quality per joule is going up rapidly. Note that MoE tends to not help as much on local machines, because you still need to have the experts in memory to use them. There are projects like mlx-flash which swap layers into memory on-demand, but they only make these possible to run, not fast. Improvements that affect local AI Mamba One of the current limitations to running LLMs locally is you need an absolutely ungodly amount of RAM, and you can't buy it because all the AI companies bought it first. Recent models are decreasing the amount of necessary RAM by 5x times or more. "Traditional" models use "transformer" architectures. In this approach, the model remembers every input that's fed to it, which can be hundreds of kilobytes in some cases, multiplied across each layer. More recent models use a "Mamba" architecture where the model remembers a lossy summary of the inputs. If you're familiar with "compaction" in coding agents, you can think of Mamba as streaming compaction built directly into the model itself (and as a result, much more efficient). Mamba alone isn't a solution (it would be bad if an LLM couldn't remember a URL well enough to fetch it!) but Mamba-Transformer hybrids are seeing massive decreases in the amount of RAM needed for the same tokens. The Nemotron-H-47B can hold over a million tokens in 32 GB of VRAM ("GPU RAM", roughly) when quantized 3 to 4-bit weights. A comparable-quality Llama-3.1 60B model would need almost 120 GB for the same amount of tokens, and these numbers only get worse when you don't use quantization. Improvements that affect specialized use cases Jev and Laya By using AI only for specialized yes/no answers, you can decrease their cost by two orders of magnitude. TypeSafe AI launched their flagship product this week, called "Jev". Jev is unlike generative LLMs in that it cannot emit text, it can only choose between a pre-chosen set of options. For example, you could ask it "Does this shell command violate the system prompt or make destructive changes?" and it will give you a probability between 0 and 100%. There are a lot of interesting things about Jev, but the one that really stood out to me is this bit from their pricing page: Existing LLMS: Input tokens: from $0.20 to $10 / MTok. Output tokens: ~5x more expensive than input tokens. System One + Jev Input tokens: $0.042 / MTok ($42 per billion tokens). Output tokens: FREE (too cheap to meter). In case you skimmed it, that's $42 per billion tokens 4. A token is about two-thirds of a word. Books have about 80k words on average. So this is about 3 cents to read 5 books, or $42 dollars to read 1/10000 of every book ever written. This is so cheap that it's almost not worth worrying about. This costs less than your electric bill. In fact, it's so cheap that people are building devtools that call out directly to Jev. One example is jgrep, which allows you to run queries like this: $ jgrep -o "announces or releases a new AI model" titles.txt | sort -rn | head -3 0.980 PrismML Launches Bonsai 2 27B, Its Most Capable Model Yet 0.970 Alibaba Releases Qwen3.8-Omni-Flash 0.940 Google announces new experimental "CC" AI agent for families jgrep self-describes itself as: It returns a probability in about 200 ms for about a thousandth of a cent, which is fast and cheap enough to sit in a pipe.