Open-weight models surge to 29% of volume, price per token flattens

Peter Bubenik · Vercel · · Source
Open-weight models surge to 29% of volume, price per token flattens

Concept 1: What is an AI Gateway?

The Basic Idea

An AI Gateway is a routing layer that sits between applications and AI providers.

Your App → AI Gateway → [Anthropic / OpenAI / Google / DeepSeek / etc.]

Why It Matters

  • It routes tens of trillions of tokens monthly
  • It gives a real-world view of how enterprises actually use AI
  • It tracks volume, spend, and pricing across all providers simultaneously

Key Insight

Because the gateway sees all traffic, it can reveal patterns that no single AI lab could see on its own — like which models get used for which tasks.


Concept 2: Tokens, Volume, and Spend — and Why They Differ

What Are Tokens?

Tokens are the unit of measurement for AI text processing. Roughly:

  • 1 token ≈ ¾ of a word
  • Every request consumes tokens (input + output)

Volume vs. Spend

These two metrics don't move together, and that gap is the most important signal in this report.

MetricJune Growth
Token Volume+29%
Spend+27%
Price per Token~Flat

Why the Gap Exists

High Volume + Low Price = Open-weight models (e.g., DeepSeek)
Low Volume  + High Price = Frontier models (e.g., Anthropic Claude)

Example from the data:

  • DeepSeek = 22.6% of tokens but a tiny fraction of spend
  • Anthropic = 32% of tokens but 61% of spend

Key Insight

Volume tells you how much AI is being used. Spend tells you what kind of AI work is being done. They measure different things.


Concept 3: Open-Weight vs. Closed-Weight (Frontier) Models

Definitions

Open-weight models:

  • The model weights (the trained parameters) are publicly released
  • Anyone can download, run, and modify them
  • Examples: DeepSeek, GLM 5.2, Qwen, Gemma
  • Typically much cheaper to run

Closed-weight (frontier) models:

  • Weights are proprietary — you access them only via API
  • Examples: Claude (Anthropic), GPT-4 (OpenAI), Gemini (Google)
  • Typically more expensive but more capable on complex tasks

The June 2026 Shift

April 2026:  Open-weight = 11% of tokens
June 2026:   Open-weight = 29% of tokens

That's nearly tripled in two months — but on only 4% of spend.

Why Companies Use Both

This is the core strategic insight:

Routine/High-Volume Work → Open-weight (cheap, fast)
High-Stakes/Complex Work → Frontier (expensive, reliable)

Concept 4: Routing Discipline — The Strategic Balancing Act

What Is Routing Discipline?

It's the deliberate practice of matching the right model to the right task based on cost and risk.

How It Works in Practice

Task TypeRisk LevelModel ChoiceWhy
Summarizing documentsLowOpen-weight (DeepSeek)Cheap, good enough
Coding agentsHighAnthropic ClaudeMistakes are costly
Back-office agentsHighAnthropic ClaudeErrors have real consequences
App generationHighAnthropic ClaudeQuality critical

The Price-Flattening Effect

This routing discipline explains why the average price per token stayed flat in June, despite two opposing forces:

↓ Pressure: More cheap open-weight volume being added
↑ Pressure: Frontier model prices rose ~12% per token

Net result: Flat average price per token

These two forces offset each other — not by accident, but because companies are strategically routing work.


Concept 5: Market Concentration — Who Dominates Where?

Text Generation

Anthropic dominates spend (high-value work):

Anthropic:  61% of spend, 32% of tokens
OpenAI:     16% of spend, 10% of tokens
DeepSeek:   ~22% of tokens, negligible spend

A Telling OpenAI Signal

OpenAI's token share fell (12.5% → 10.3%) while its spend share rose (13.3% → 16.1%).

This means customers sent OpenAI less work, but harder work — pushing its effective cost per token up ~50% relative to the market.

Image Generation

A two-lab race:

OpenAI GPT Image:    53% of images, 52% of spend
Google Nano Banana:  39% of images, 43% of spend
Everyone else:       <5%

Video Generation

A different dynamic — Chinese labs dominate spend:

ByteDance Seedance:  49% of spend, 34% of videos (premium positioning)
xAI Grok Imagine:   42% of videos, 19% of spend (volume/discount positioning)
Chinese labs total: ~67% of video spend

Key Insight: Modality Leadership

Text:         Anthropic (spend leader)
Images:       OpenAI (volume + spend leader)
Video volume: xAI
Video spend:  ByteDance

To use the best model in every modality, you need to route across at least 3 different labs — which is exactly why an AI Gateway exists.


Concept 6: Model Adoption Velocity

How Fast Do New Models Get Adopted?

The article gives two examples that show adoption is accelerating:

Claude Fable 5 (June 9 release):

Day 1 → Day 4: Reached 22% of Opus 4.8's request volume

Then suspended by US export controls on June 12.

GLM 5.2 (June 16 release):

2 weeks after launch: 50x daily token volume growth
                      Ranked #11 overall, #7 on peak days
                      Captured 76% of its model family's June tokens

Compare that to the previous fastest adoption:

Gemini 3.1 Pro: Took until its SECOND MONTH to reach similar family share
GLM 5.2:        Did it in TWO WEEKS

Why Adoption Is Accelerating

  • AI Gateways make switching easy — one integration, many models
  • Open-weight models with MIT licenses remove legal friction
  • Competitive pricing (~1/5th of Opus 4.8) creates strong incentive to trial

Concept 7: Export Controls as a Market Force

What Happened with Claude Fable 5?

June 9:  Fable 5 released → rapid adoption begins
June 12: US export-control directive takes effect
         Anthropic suspends access to comply
June 30: Controls lifted
July 1:  Access resumes

Why This Matters as a Concept

Export controls are a non-market force that can instantly reshape AI usage patterns. This is a new risk category for enterprises:

A model you depend on can become unavailable overnight due to geopolitical/regulatory action — not technical failure.

This reinforces the value of multi-model routing — if one model goes offline, traffic can be redirected.


Concept 8: B2B vs. B2C Usage Patterns

The Split

SegmentToken VolumeSpend
B2B46%60%
B2C43%26%

What This Tells Us

  • B2B work costs more per token — businesses use frontier models for high-stakes tasks
  • B2C work is cheaper per token — consumer apps optimize for cost at scale

Google's Interesting Position

Google's token share:
- Personal assistant tasks: 57%
- Education tasks:          54%
- Coding agent tasks:       <2%

Google dominates consumer-shaped workloads but is nearly absent from high-stakes enterprise workloads — which explains why it has high token volume but low spend share.


Summary: The Big Picture

All these concepts connect into one coherent story:

1. AI usage is growing fast (29% MoM volume)

2. But companies are getting smarter about HOW they spend:
   → Cheap open-weight models for routine work
   → Expensive frontier models for critical work

3. This "routing discipline" keeps average prices flat
   even as both cheap and expensive options grow

4. No single lab wins everywhere:
   → Different leaders in text, image, and video
   → Multi-lab routing is now a necessity, not a luxury

5. New risks are emerging:
   → Export controls can suspend models overnight
   → Adoption velocity is accelerating (more disruption, faster)

The core lesson: Enterprise AI strategy in 2026 is about intelligent routing across a diverse model portfolio, not loyalty to a single provider.

More to study