Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

Peter Bubenik · Vercel · · Source
Image for Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

Step-by-Step Teaching

Step 1: Understand the Two Types of AI Models

Before reading any market data, you need this foundation.

TypeDefinitionExample
Closed-weightProprietary models, weights not publicly availableGPT-6 Astra, Claude Fable 5
Open-weightWeights publicly released, can be self-hostedModels teams can run independently

Why it matters: Open-weight models typically cost less because teams can run them on their own infrastructure, bypassing per-token fees charged by labs.


Step 2: Understand What "Token Volume" vs "Spend" Actually Measures

These are two different metrics that tell different stories.

Token Volume = HOW MUCH is being processed
Spend        = HOW MUCH MONEY is being paid

Critical insight: A model can have HIGH token volume but LOW spend if it's cheap, and vice versa.

Real example from the data:

Luna and Nano processed twice as many tokens as Astra and Sol, but generated only one-ninth the spending

This means cheap models dominate usage, but expensive models dominate revenue.


Step 3: Learn the Core Market Shift — Open-Weight Takeover

The Timeline

December 2025:  Open-weight = 7% of tokens
August 2026:    Open-weight = 56% of tokens
                ↑
         Majority for the FIRST TIME

Why Did This Happen?

Three reinforcing forces:

1. Capability improved → Open-weight models became good enough for production workloads

2. Price dropped → Average token cost fell 23.2% in August alone (third consecutive monthly drop)

3. Budget optimization → Teams learned to use frontier models only when necessary

The resulting strategy enterprises adopted:

Cheap/open-weight models → Routine tasks (high volume)
Frontier models          → Tasks that justify premium (low volume)

Step 4: Understand the "Good Enough" Principle in Model Selection

This is one of the most important concepts in the article.

"Production workloads that justify a frontier model don't always need the most expensive one. They need one that's good enough."

The Fable → Opus 5 Case Study

ModelPriceShare BeforeShare After
Fable 5 (most capable)2x price13.2%4.9%
Opus 5 (tier below)1x priceLow22.5%

What happened:

  • Fable 5 access was restored July 1 → spend surged
  • Opus 5 launched end of July at half the price
  • 9 out of 10 teams using Fable cut their usage
  • Most moved to Opus 5, not a competitor

The lesson: When a cheaper model handles the same workload, teams will downgrade within the same lab rather than pay for capability they don't need.


Step 5: Understand Lab Loyalty vs. Model Loyalty

This is a nuanced but critical distinction.

The Key Principle

Loyalty follows MODEL PROFILE, not BRAND

What does "model profile" mean? The combination of:

  • Price point
  • Capability level
  • Task fit

Two Contrasting Examples

✅ Anthropic succeeded at retention:

Teams left Fable 5
    ↓
Moved to Opus 5 (same lab, lower tier)
    ↓
Anthropic retained 64% of all gateway spend

Anthropic held the top two spend positions every month since December, even as the specific models changed.

❌ Google failed at retention:

Teams left Gemini 3 Flash
    ↓
Moved to OpenAI, Anthropic, DeepSeek
    ↓
Google's token volume share: 30% → 5%
Gemini 3 Flash caused 22 of 25 percentage points lost

Why the Different Outcomes?

FactorAnthropicGoogle
Replacement model available?Yes (Opus 5)No adequate replacement
Price advantage offered?Yes (half price)No
Capability maintained?YesNo relative advantage

Step 6: Analyze a Model Launch — The Astra Case Study

Setup

  • Fable 5.1 launched September 1 (Anthropic)
  • GPT-6 Astra launched September 3 (OpenAI)
  • Both priced identically
  • Launched just 2 days apart

Results After 12 Days

Astra:     7.7% of ALL gateway spend
Fable 5.1: 3.7% of ALL gateway spend
           ↑
      Astra = 2x Fable 5.1

Astra also captured 1 in 3 OpenAI dollars within 48 hours of launch.

What This Tells Us

  1. Brand momentum matters at launch — OpenAI's existing user base adopted Astra rapidly
  2. Same price = direct competition — When price is equal, capability perception and trust decide
  3. Portfolio strategy works — OpenAI uses cheap models (Luna, Nano) for volume AND premium models (Astra) for high-value spend

Step 7: Synthesize — The Framework for Understanding AI Market Dynamics

Now you can build a complete mental model:

┌─────────────────────────────────────────────────────┐
│           AI MODEL MARKET DYNAMICS FRAMEWORK         │
├─────────────────────────────────────────────────────┤
│                                                      │
│  PRICE DROPS → More open-weight adoption             │
│      ↓                                               │
│  Teams optimize: cheap models for volume,            │
│  frontier models for premium tasks                   │
│      ↓                                               │
│  Labs compete on: capability + price fit             │
│      ↓                                               │
│  Retention = having the RIGHT model at each tier     │
│      ↓                                               │
│  Labs without tier coverage LOSE customers           │
│  to competitors entirely                             │
│                                                      │
└─────────────────────────────────────────────────────┘

Step 8: Test Your Understanding

Answer these questions to confirm mastery:

Q1: Why did Anthropic retain 64% of spend even though teams abandoned its most expensive model?

Q2: A lab releases a new model at the same price as its predecessor with no capability improvement. Based on this article, what would you predict happens to its market share?

Q3: Why does high token volume NOT necessarily mean high revenue for a lab?

Q4: What two things must a replacement model offer to retain customers within the same lab?


Answers

A1: Because Anthropic had Opus 5 at half the price — teams stepped down within Anthropic's lineup rather than switching labs.

A2: It would likely lose share, as Google's Gemini 3 Flash demonstrated — no relative advantage on capability or price means customers move to better-fit models elsewhere.

A3: Because cheap/open-weight models generate massive token volume but minimal revenue per token. Spend and volume are separate metrics.

A4: The replacement must offer a price advantage OR capability advantage (ideally both) while handling the same workloads as the model being replaced.

More to study