Before reading any market data, you need to understand the units of measurement.
| Term | Definition | Simple Analogy |
|---|---|---|
| Token | The basic unit of AI text processing (roughly ΒΎ of a word) | Like a "unit" of electricity |
| Token Volume | Total number of tokens processed | Like kilowatt-hours consumed |
| Spend | Total money paid, calculated at published list prices | Your electricity bill |
| Price per Token | Total spend Γ· total volume | Cost per kilowatt-hour |
| AI Gateway | A routing layer that directs requests from applications to AI labs | Like a telephone exchange |
| Closed-weight models | AI models where the underlying code is proprietary (e.g., Claude, GPT) | A secret recipe |
| Open-weight models | AI models where weights are publicly available (e.g., DeepSeek) | An open-source recipe |
If volume grows 59% but spend only grows 37%, what must be happening to price per token? Answer: Price per token must be falling, because you are buying more but paying proportionally less.
This is the most important concept in the article.
Price per Token = Total Spend Γ· Total Volume
Volume grew: +59%
Spend grew: +37%
Result: Price per token FELL 13.6%
Think of it like buying fruit at a market:
In AI terms: companies routed more traffic to cheaper models, pulling the average price down even though total spending increased.
Volume and spend can both grow while price per token falls. This happens when the mix of models shifts toward cheaper options.
BY TOKEN VOLUME BY SPEND (REVENUE)
βββββββββββββββββ ββββββββββββββββββ
1. Anthropic ~30% 1. Anthropic 65%
2. DeepSeek ~25% 2. OpenAI ~12%
3. OpenAI ~13% 3. Google smaller
4. Google ~11% 4. Open-weight ~9%
Anthropic has 30% of volume but 65% of spend. DeepSeek has 25% of volume but a tiny share of spend.
This tells you something critical: not all tokens are equal in value.
| Model | Relative Price |
|---|---|
| Anthropic average | 4.4Γ gateway average |
| OpenAI GPT-5-Nano | 1/6 of gateway average |
| DeepSeek V4 Flash | 1/16 of gateway average |
Buyers sort by task first, then by price within that tier. DeepSeek and Claude Opus are not competing with each other β they serve different jobs.
April 2026: DeepSeek = <1% of gateway tokens
Google = ~40% of gateway tokens
July 2026: DeepSeek = ~25% of gateway tokens
Google = ~11% of gateway tokens
Three factors combined:
Winning on volume β winning on revenue. DeepSeek became #2 by tokens while remaining a minor player by dollars.
The article explains this with a simple logic chain:
Cheap models (DeepSeek, GPT-5-Nano)
β Can handle simple/medium tasks
β Cannot handle complex agent work
Anthropic models
β Handle complex tasks (coding agents, etc.)
β Collect 80%+ of coding agent spend
Therefore:
β Customers who need complex work MUST pay Anthropic's price
β Customers who can use cheap models ALREADY left
When Claude Fable 5 came back after a 3-week suspension:
This shows: the work exists regardless of who does it. Customers who needed that capability found it immediately when it returned.
Open-weight volume share: 11% β 29% (tripling)
Open-weight spend share: stayed under 4 cents per dollar
This means: open-weight models were getting popular but not making money.
Open-weight volume share: 36% (continued growth)
Open-weight spend share: ~9 cents per dollar (MORE THAN DOUBLED)
Two new models: Kimi K3 (Moonshot) and GLM 5.2 (Z.ai)
These are the first open-weight models that:
Open-weight models previously competed only on price. Kimi K3 and GLM 5.2 are the first to compete on capability, entering the premium tier of work.
| Lab | Volume Share | Spend Share |
|---|---|---|
| Google (Nano Banana) | 45% | ~50% |
| OpenAI (GPT Image) | 42% | ~50% |
Key model: Gemini 3.1 Flash Lite Image drove nearly all of Google's gain.
Interesting note: Google took the image volume lead in the same month it fell to 4th place in text token volume.
| Lab | Position |
|---|---|
| ByteDance (Seedance) | #1 in both volume AND spend |
| xAI (Grok Imagine) | #2 (was #1 in June) |
| Chinese labs combined | ~70% of video spend |
TREND 1: Commoditization at the bottom
βββββββββββββββββββββββββββββββββββββ
Cheap models get cheaper and more popular
β Price per token falls
β Volume explodes
β Revenue stays concentrated at top
TREND 2: Premium consolidation at the top
βββββββββββββββββββββββββββββββββββββββββ
Anthropic holds complex workloads
β Price premium widens (4.4Γ vs 3.4Γ in June)
β Revenue share stays at 65%
β No cheap alternative for complex tasks
TREND 3: Open-weight models moving upmarket
ββββββββββββββββββββββββββββββββββββββββββββ
Kimi K3 and GLM 5.2 enter agent work
β First open-weight revenue breakthrough
β Open-weight spend share doubles
β Frontier labs' combined share falls below 90%
HIGH CAPABILITY / HIGH PRICE
β
β Anthropic (Claude Opus, Fable 5)
β β Complex agents, coding
β
β OpenAI (GPT-5 family)
β β Mixed workloads
β
β Kimi K3, GLM 5.2 β NEW ENTRANTS HERE
β β Long-horizon agent work
β
β Google, DeepSeek
β β Consumer-facing, simple tasks
β
LOW CAPABILITY / LOW PRICE
A lab can dominate token volume while capturing almost no revenue if its prices are low enough. DeepSeek is the clearest example.
Enterprise buyers choose the tier (premium vs. cheap) based on what the task requires, then optimize within that tier on price. This is why price competition is "happening in the part of the market with the least money in it."
Average price per token is driven more by which models companies choose than by individual model price changes. 81% of July's tokens ran on models that didn't exist 6 months ago.
Anthropic's pricing power comes not from technical lock-in but from the gap between what cheap models can do and what its customers need. When that gap closes, so does the premium.
Test your understanding with these questions:
Basic: If volume grows 59% and spend grows 37%, by approximately how much does price per token change? (Hint: use the formula)
Intermediate: Why did DeepSeek gaining 25% of token volume not significantly hurt Anthropic's 65% revenue share?
Intermediate: What made Kimi K3 different from previous open-weight models like DeepSeek in terms of market impact?
Advanced: The article says "price competition is happening in the part of the market with the least money in it." Explain what this means and why it is true based on the data.
Advanced: If you were advising a startup building an AI-powered coding assistant, which lab would the data suggest you use, and why? What would change your answer if you were building a consumer chatbot instead?
Price per token fell approximately 13.6% β volume grew faster than spend, so the average cost per unit dropped.
DeepSeek's tokens are extremely cheap (1/16th of average price), so even massive volume generates little revenue. Anthropic's customers need capabilities that DeepSeek cannot provide, so they don't switch.
Previous open-weight models (like DeepSeek) competed only on price for simple tasks. Kimi K3 entered agent work β complex, multi-step tasks β at 11Γ DeepSeek's price, capturing revenue that previously only went to closed-weight labs.
The cheap tier (open-weight, DeepSeek) has most of the volume but almost none of the money. Price wars there don't affect the majority of AI spending, which stays concentrated in Anthropic's premium tier.
Coding assistant: Use Anthropic β it collects 80%+ of coding agent spend, suggesting it performs best for this task. Consumer chatbot: Consider DeepSeek or Google β high volume, low cost, sufficient for simpler consumer interactions.