
Before reading any market data, you need this foundation.
| Type | Definition | Example |
|---|---|---|
| Closed-weight | Proprietary models, weights not publicly available | GPT-6 Astra, Claude Fable 5 |
| Open-weight | Weights publicly released, can be self-hosted | Models teams can run independently |
Why it matters: Open-weight models typically cost less because teams can run them on their own infrastructure, bypassing per-token fees charged by labs.
These are two different metrics that tell different stories.
Token Volume = HOW MUCH is being processed
Spend = HOW MUCH MONEY is being paid
Critical insight: A model can have HIGH token volume but LOW spend if it's cheap, and vice versa.
Real example from the data:
Luna and Nano processed twice as many tokens as Astra and Sol, but generated only one-ninth the spending
This means cheap models dominate usage, but expensive models dominate revenue.
December 2025: Open-weight = 7% of tokens
August 2026: Open-weight = 56% of tokens
↑
Majority for the FIRST TIME
Three reinforcing forces:
1. Capability improved → Open-weight models became good enough for production workloads
2. Price dropped → Average token cost fell 23.2% in August alone (third consecutive monthly drop)
3. Budget optimization → Teams learned to use frontier models only when necessary
The resulting strategy enterprises adopted:
Cheap/open-weight models → Routine tasks (high volume)
Frontier models → Tasks that justify premium (low volume)
This is one of the most important concepts in the article.
"Production workloads that justify a frontier model don't always need the most expensive one. They need one that's good enough."
| Model | Price | Share Before | Share After |
|---|---|---|---|
| Fable 5 (most capable) | 2x price | 13.2% | 4.9% |
| Opus 5 (tier below) | 1x price | Low | 22.5% |
What happened:
The lesson: When a cheaper model handles the same workload, teams will downgrade within the same lab rather than pay for capability they don't need.
This is a nuanced but critical distinction.
Loyalty follows MODEL PROFILE, not BRAND
What does "model profile" mean? The combination of:
✅ Anthropic succeeded at retention:
Teams left Fable 5
↓
Moved to Opus 5 (same lab, lower tier)
↓
Anthropic retained 64% of all gateway spend
Anthropic held the top two spend positions every month since December, even as the specific models changed.
❌ Google failed at retention:
Teams left Gemini 3 Flash
↓
Moved to OpenAI, Anthropic, DeepSeek
↓
Google's token volume share: 30% → 5%
Gemini 3 Flash caused 22 of 25 percentage points lost
| Factor | Anthropic | |
|---|---|---|
| Replacement model available? | Yes (Opus 5) | No adequate replacement |
| Price advantage offered? | Yes (half price) | No |
| Capability maintained? | Yes | No relative advantage |
Astra: 7.7% of ALL gateway spend
Fable 5.1: 3.7% of ALL gateway spend
↑
Astra = 2x Fable 5.1
Astra also captured 1 in 3 OpenAI dollars within 48 hours of launch.
Now you can build a complete mental model:
┌─────────────────────────────────────────────────────┐
│ AI MODEL MARKET DYNAMICS FRAMEWORK │
├─────────────────────────────────────────────────────┤
│ │
│ PRICE DROPS → More open-weight adoption │
│ ↓ │
│ Teams optimize: cheap models for volume, │
│ frontier models for premium tasks │
│ ↓ │
│ Labs compete on: capability + price fit │
│ ↓ │
│ Retention = having the RIGHT model at each tier │
│ ↓ │
│ Labs without tier coverage LOSE customers │
│ to competitors entirely │
│ │
└─────────────────────────────────────────────────────┘
Answer these questions to confirm mastery:
Q1: Why did Anthropic retain 64% of spend even though teams abandoned its most expensive model?
Q2: A lab releases a new model at the same price as its predecessor with no capability improvement. Based on this article, what would you predict happens to its market share?
Q3: Why does high token volume NOT necessarily mean high revenue for a lab?
Q4: What two things must a replacement model offer to retain customers within the same lab?
A1: Because Anthropic had Opus 5 at half the price — teams stepped down within Anthropic's lineup rather than switching labs.
A2: It would likely lose share, as Google's Gemini 3 Flash demonstrated — no relative advantage on capability or price means customers move to better-fit models elsewhere.
A3: Because cheap/open-weight models generate massive token volume but minimal revenue per token. Spend and volume are separate metrics.
A4: The replacement must offer a price advantage OR capability advantage (ideally both) while handling the same workloads as the model being replaced.