
An AI Gateway is a routing layer that sits between applications and AI providers.
Your App → AI Gateway → [Anthropic / OpenAI / Google / DeepSeek / etc.]
Because the gateway sees all traffic, it can reveal patterns that no single AI lab could see on its own — like which models get used for which tasks.
Tokens are the unit of measurement for AI text processing. Roughly:
These two metrics don't move together, and that gap is the most important signal in this report.
| Metric | June Growth |
|---|---|
| Token Volume | +29% |
| Spend | +27% |
| Price per Token | ~Flat |
High Volume + Low Price = Open-weight models (e.g., DeepSeek)
Low Volume + High Price = Frontier models (e.g., Anthropic Claude)
Example from the data:
Volume tells you how much AI is being used. Spend tells you what kind of AI work is being done. They measure different things.
Open-weight models:
Closed-weight (frontier) models:
April 2026: Open-weight = 11% of tokens
June 2026: Open-weight = 29% of tokens
That's nearly tripled in two months — but on only 4% of spend.
This is the core strategic insight:
Routine/High-Volume Work → Open-weight (cheap, fast)
High-Stakes/Complex Work → Frontier (expensive, reliable)
It's the deliberate practice of matching the right model to the right task based on cost and risk.
| Task Type | Risk Level | Model Choice | Why |
|---|---|---|---|
| Summarizing documents | Low | Open-weight (DeepSeek) | Cheap, good enough |
| Coding agents | High | Anthropic Claude | Mistakes are costly |
| Back-office agents | High | Anthropic Claude | Errors have real consequences |
| App generation | High | Anthropic Claude | Quality critical |
This routing discipline explains why the average price per token stayed flat in June, despite two opposing forces:
↓ Pressure: More cheap open-weight volume being added
↑ Pressure: Frontier model prices rose ~12% per token
Net result: Flat average price per token
These two forces offset each other — not by accident, but because companies are strategically routing work.
Anthropic dominates spend (high-value work):
Anthropic: 61% of spend, 32% of tokens
OpenAI: 16% of spend, 10% of tokens
DeepSeek: ~22% of tokens, negligible spend
OpenAI's token share fell (12.5% → 10.3%) while its spend share rose (13.3% → 16.1%).
This means customers sent OpenAI less work, but harder work — pushing its effective cost per token up ~50% relative to the market.
A two-lab race:
OpenAI GPT Image: 53% of images, 52% of spend
Google Nano Banana: 39% of images, 43% of spend
Everyone else: <5%
A different dynamic — Chinese labs dominate spend:
ByteDance Seedance: 49% of spend, 34% of videos (premium positioning)
xAI Grok Imagine: 42% of videos, 19% of spend (volume/discount positioning)
Chinese labs total: ~67% of video spend
Text: Anthropic (spend leader)
Images: OpenAI (volume + spend leader)
Video volume: xAI
Video spend: ByteDance
To use the best model in every modality, you need to route across at least 3 different labs — which is exactly why an AI Gateway exists.
The article gives two examples that show adoption is accelerating:
Claude Fable 5 (June 9 release):
Day 1 → Day 4: Reached 22% of Opus 4.8's request volume
Then suspended by US export controls on June 12.
GLM 5.2 (June 16 release):
2 weeks after launch: 50x daily token volume growth
Ranked #11 overall, #7 on peak days
Captured 76% of its model family's June tokens
Compare that to the previous fastest adoption:
Gemini 3.1 Pro: Took until its SECOND MONTH to reach similar family share
GLM 5.2: Did it in TWO WEEKS
June 9: Fable 5 released → rapid adoption begins
June 12: US export-control directive takes effect
Anthropic suspends access to comply
June 30: Controls lifted
July 1: Access resumes
Export controls are a non-market force that can instantly reshape AI usage patterns. This is a new risk category for enterprises:
A model you depend on can become unavailable overnight due to geopolitical/regulatory action — not technical failure.
This reinforces the value of multi-model routing — if one model goes offline, traffic can be redirected.
| Segment | Token Volume | Spend |
|---|---|---|
| B2B | 46% | 60% |
| B2C | 43% | 26% |
Google's token share:
- Personal assistant tasks: 57%
- Education tasks: 54%
- Coding agent tasks: <2%
Google dominates consumer-shaped workloads but is nearly absent from high-stakes enterprise workloads — which explains why it has high token volume but low spend share.
All these concepts connect into one coherent story:
1. AI usage is growing fast (29% MoM volume)
2. But companies are getting smarter about HOW they spend:
→ Cheap open-weight models for routine work
→ Expensive frontier models for critical work
3. This "routing discipline" keeps average prices flat
even as both cheap and expensive options grow
4. No single lab wins everywhere:
→ Different leaders in text, image, and video
→ Multi-lab routing is now a necessity, not a luxury
5. New risks are emerging:
→ Export controls can suspend models overnight
→ Adoption velocity is accelerating (more disruption, faster)
The core lesson: Enterprise AI strategy in 2026 is about intelligent routing across a diverse model portfolio, not loyalty to a single provider.