Skip to content
HN On Hacker News ↗

Anthropic Subscriptions Offer 5x+ More Value Than OpenAI

▲ 82 points • 94 comments • by nsoonhui • 3d ago • HN discussion ↗

Pangram verdict · v3.3

We believe this text is mainly human-written, with some AI content.

7 %

AI likelihood · overall

Human
95% human-written 5% AI-generated
SEGMENTS · HUMAN 1 of 4
SEGMENTS · AI 1 of 4
WORD COUNT 1,127
PEAK AI % 74% · §2
Analyzed
Oct 6
backend: pangram/v3.3
Segments scanned
4 windows
avg 282 words each
Distribution
95 / 5%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,127 words · 4 segments analyzed

Human AI-generated
§1 Human · 4%

Subscription plans are still the primary way consumers and small businesses pay for AI. These plans are highly subsidized—as we previously explained in June—but can still make economic sense as powerful customer acquisition and marketing tools. For example, the goodwill engendered by OpenAI’s generous resets is partially responsible for the recent surge in Codex adoption and has forced Anthropic to repeatedly walk back planned subscription nerfs to avoid getting clobbered in the court of public opinion.Furthermore, because subscription plans are so heavily subsidized, they can also have a large impact on margins and revenue per MW despite only being a small portion of total revenue. Consider the following rough numbers for Anthropic.Despite being just 10% of overall revenue, subscriptions can take up over 40% of inference compute and lower blended revenue per MW by ~$36M. Subscriptions are even more important for OpenAI, as they make up a larger portion of their total revenue. For specific numbers, see our Tokenomics Model.In other words, if you want to accurately model AI lab financials, you need to understand their subscription limits.So how do limits actually work? The way to think about subscriptions is that your monthly payment grants you some numbers of “credits”. Each (model, token type) combo consumes a different amount of credits. Because credit cost ratios can differ dramatically from API price ratios, the “value” of the same plan changes depending on what model and workload you’re running. Put differently, it doesn’t make sense to say a plan is worth $X in isolation—you need to consider the full (plan, model, workload) tuple.Here’s how the API-equivalent value of the same $200/month Claude plan changes depending on the model and workload.A single static snapshot isn’t enough either. Labs publicly change their limits all the time with promos and new model releases. They can also silently change limits whenever they want by tweaking credit costs. Ideally, you’d re-check the cost of each (plan, model, token type) daily so you can surface any changes in real-time.This is exactly what SemiAnalysis has done with our new Subscriptions Dashboard available exclusively to Tokenomics Model subscribers. Besides every OpenAI and Anthropic subscription, we also track Meta, SpaceXAI, Cursor, Cognition, Z.ai, MiniMax, and Moonshot.A screenshot of our dashboard showing a small subset of the available data. Source: SemiAnalysis Tokenomics ModelNew providers, plans, and models will be added as soon as they’re released. The rest of the article will give an overview of everyone’s current limits. All token and dollar amounts below assume agentic usage unless otherwise specified.Tokens are generally priced per MTok (million tokens) across the following types:Input: Fresh tokens added to the LLM’s context window that aren’t cachedCache write: Input tokens that are cached for multi-turn conversations. Generally slightly more expensive than regular input tokens.Cache read: Tokens from previous turns that are already cached in the conversation. Generally extremely cheap compared to uncached tokens.Output: Tokens generated by the model. The most expensive token type.Additionally, some models charge more for tokens above a certain context window. Those models typically compact before that expensive context window is reached, so we also keep our measurements below it.Source: OpenAISubscription plans don’t expose this fine grained pricing. Instead, they provide a simple 0 to 100% usage meter across 5-hour and 7-day windows, and sometimes a separate meter for a model like Fable. Subscription tiers are then differentiated by usage multipliers relative only to the provider’s other plans. OpenAI for example used to advertise “Expanded Codex usage” in their $20 Plus plan, “5x more usage than Plus” in the $100 Pro plan, and “20x more usage” in the $200 Pro plan until they cut their $200 plan usage in half and removed all relative usage from their pricing page.Screenshot of Claude Code meters. Source: Claude Code CLIWe compute the subscription-rate of each (plan, model, token type) triple by running experiments that isolate one token type at a time and watching how far the model provider’s meter moves. An experiment is some number of repeated calls using a specific prompt that maximizes one token type while minimizing all the others.Input, cache writes, and cache reads share the same prompt template. We use a portion of War and Peace since some models will refuse to respond if given large blocks of gibberish. For the input token experiments, we use a random tag on every call to ensure nothing gets cached. The cache write experiments run the same way but with the prompt marked for caching, so each new tag forces a new cache entry. The cache read experiments use a fixed tag, so the first call writes the cache and every repeat reads it.For the output token experiments, we used a technical essay to force long outputs because models will refuse mechanical prompts like “repeat SemiAnalysis 100,000 times”.For every call we record two things: how many tokens of each type the provider billed, and what the usage meter read. Turning those into a price takes three steps.Providers generally report usage on a meter that moves in fixed amounts. This could be a whole percentage point, a credit, a cent etc... A single request often does not move the meter at all, so the cost of one request cannot be measured directly.Instead, we measure in steps.

§2 AI · 74%

As requests run, we keep a running total of the tokens used. Each time the meter rises, we store that total. The tokens used between two meter moves make up one step, which is the cost of moving the meter by one unit.We drop two partial steps. When a run starts, the meter is already part of the way to its next move, so the tokens before the first move are less than a full step and would make the rate look higher than it is. The tokens after the last move never finish a step, so we drop those as well.To get a rate, we add up how far the meter moved over the complete steps and divide by the tokens used in them.

§3 Mixed · 48%

Providers charge differently for input, output, and cached tokens, so we calculate a separate rate for each.Caveats:Because we see the meter only once per request, a single step can be off by up to one request. These errors don’t add up with more consecutive steps, so we track a range instead of a single number, and the range gets narrower the more steps we count. We keep adding steps until the range is within ±5%, then report the rate across all of them.No request contains only one type of token.

§4 Mixed · 62%

Some plans, for example, require a set of instructions on every request. We subtract that extra amount using the prices we measured for the other types.Reading from the cache costs almost nothing on some plans and may not move the meter at all.