AI Model Usage Shifts: DeepSeek Rises, Costs Fall

Peter Bubenik Β· Vercel Β· Β· Source
Image for DeepSeek overtakes Google on volume, cost per token falls 13.6%

Step-by-Step Study Material

Step 1: Understanding the Basic Vocabulary

Before reading any market data, you need to understand the units of measurement.

Key Terms

TermDefinitionSimple Analogy
TokenThe basic unit of AI text processing (roughly ΒΎ of a word)Like a "unit" of electricity
Token VolumeTotal number of tokens processedLike kilowatt-hours consumed
SpendTotal money paid, calculated at published list pricesYour electricity bill
Price per TokenTotal spend Γ· total volumeCost per kilowatt-hour
AI GatewayA routing layer that directs requests from applications to AI labsLike a telephone exchange
Closed-weight modelsAI models where the underlying code is proprietary (e.g., Claude, GPT)A secret recipe
Open-weight modelsAI models where weights are publicly available (e.g., DeepSeek)An open-source recipe

βœ… Check Your Understanding

If volume grows 59% but spend only grows 37%, what must be happening to price per token? Answer: Price per token must be falling, because you are buying more but paying proportionally less.


Step 2: Understanding the Core Economic Relationship

This is the most important concept in the article.

The Token Economics Formula

Price per Token = Total Spend Γ· Total Volume

July 2026 Example

Volume grew:  +59%
Spend grew:   +37%
Result:       Price per token FELL 13.6%

Why Does This Happen?

Think of it like buying fruit at a market:

  • You bought more fruit overall (volume up 59%)
  • You spent more money overall (spend up 37%)
  • But you bought cheaper fruit in your mix (more bananas, fewer strawberries)
  • So your average price per piece of fruit fell

In AI terms: companies routed more traffic to cheaper models, pulling the average price down even though total spending increased.

πŸ”‘ Key Insight

Volume and spend can both grow while price per token falls. This happens when the mix of models shifts toward cheaper options.


Step 3: The Market Structure β€” Who Are the Players?

The AI Lab Hierarchy (July 2026)

BY TOKEN VOLUME          BY SPEND (REVENUE)
─────────────────        ──────────────────
1. Anthropic  ~30%       1. Anthropic   65%
2. DeepSeek   ~25%       2. OpenAI      ~12%
3. OpenAI     ~13%       3. Google      smaller
4. Google     ~11%       4. Open-weight ~9%

Notice the Disconnect

Anthropic has 30% of volume but 65% of spend. DeepSeek has 25% of volume but a tiny share of spend.

This tells you something critical: not all tokens are equal in value.


Step 4: Understanding the Two-Tier Market

Tier 1: Premium / Closed-Weight Models

  • Examples: Anthropic Claude, OpenAI GPT-5
  • Characteristics: Expensive, high capability, used for complex tasks
  • Use case example: Coding agents, complex reasoning

Tier 2: Cheap / Open-Weight Models

  • Examples: DeepSeek V4 Flash, Kimi K3, GLM 5.2
  • Characteristics: Very cheap, high volume, used for simpler tasks
  • Use case example: Consumer-facing assistants, personal queries

Price Comparison Table

ModelRelative Price
Anthropic average4.4Γ— gateway average
OpenAI GPT-5-Nano1/6 of gateway average
DeepSeek V4 Flash1/16 of gateway average

πŸ”‘ Key Insight

Buyers sort by task first, then by price within that tier. DeepSeek and Claude Opus are not competing with each other β€” they serve different jobs.


Step 5: DeepSeek's Rise β€” A Case Study in Market Disruption

The Timeline

April 2026:  DeepSeek = <1% of gateway tokens
             Google   = ~40% of gateway tokens

July 2026:   DeepSeek = ~25% of gateway tokens
             Google   = ~11% of gateway tokens

How Did This Happen?

Three factors combined:

  1. Extreme price advantage β€” DeepSeek V4 Flash costs 1/16th of the average token price
  2. Consumer-facing work shifted β€” Google's personal-assistant token share fell by more than half; most went to DeepSeek
  3. One model dominated β€” DeepSeek V4 Flash alone ran more tokens than all of Google combined

What DeepSeek Did NOT Do

  • It did not take significant revenue share (spend barely moved)
  • It did not compete with Anthropic in premium workloads
  • It did not displace closed-weight labs from high-value tasks

πŸ”‘ Key Insight

Winning on volume β‰  winning on revenue. DeepSeek became #2 by tokens while remaining a minor player by dollars.


Step 6: Anthropic's Premium Strategy β€” Why It Works

The Numbers

  • Anthropic tokens cost 4.4Γ— the average of all other labs
  • Anthropic collected 65% of all gateway spending
  • Anthropic has no model at the cheap end of the market

Why Don't Customers Leave?

The article explains this with a simple logic chain:

Cheap models (DeepSeek, GPT-5-Nano)
    β†’ Can handle simple/medium tasks
    β†’ Cannot handle complex agent work

Anthropic models
    β†’ Handle complex tasks (coding agents, etc.)
    β†’ Collect 80%+ of coding agent spend

Therefore:
    β†’ Customers who need complex work MUST pay Anthropic's price
    β†’ Customers who can use cheap models ALREADY left

The Switching Cost Paradox

  • Switching models on AI Gateway = one line of code (technically easy)
  • Yet Anthropic retains customers because the capability gap is the real barrier, not technical lock-in

Evidence: Claude Fable 5 Returned After a Ban

When Claude Fable 5 came back after a 3-week suspension:

  • Daily volume returned to exactly pre-ban levels
  • But 9 in 10 teams running it were new customers

This shows: the work exists regardless of who does it. Customers who needed that capability found it immediately when it returned.


Step 7: Open-Weight Models β€” The Spend Breakthrough

Historical Pattern (April–June 2026)

Open-weight volume share:  11% β†’ 29% (tripling)
Open-weight spend share:   stayed under 4 cents per dollar

This means: open-weight models were getting popular but not making money.

July 2026 Breakthrough

Open-weight volume share:  36% (continued growth)
Open-weight spend share:   ~9 cents per dollar (MORE THAN DOUBLED)

What Changed?

Two new models: Kimi K3 (Moonshot) and GLM 5.2 (Z.ai)

These are the first open-weight models that:

  • Run long-horizon agent work (complex, multi-step tasks)
  • Charge 11Γ— DeepSeek's rate per token
  • Compete with workloads historically owned by closed-weight labs

πŸ”‘ Key Insight

Open-weight models previously competed only on price. Kimi K3 and GLM 5.2 are the first to compete on capability, entering the premium tier of work.


Step 8: Media Models β€” Images and Video

Image Generation (July 2026)

LabVolume ShareSpend Share
Google (Nano Banana)45%~50%
OpenAI (GPT Image)42%~50%

Key model: Gemini 3.1 Flash Lite Image drove nearly all of Google's gain.

Interesting note: Google took the image volume lead in the same month it fell to 4th place in text token volume.

Video Generation (July 2026)

LabPosition
ByteDance (Seedance)#1 in both volume AND spend
xAI (Grok Imagine)#2 (was #1 in June)
Chinese labs combined~70% of video spend

Step 9: Synthesizing the Big Picture

The Three Simultaneous Trends

TREND 1: Commoditization at the bottom
─────────────────────────────────────
Cheap models get cheaper and more popular
β†’ Price per token falls
β†’ Volume explodes
β†’ Revenue stays concentrated at top

TREND 2: Premium consolidation at the top
─────────────────────────────────────────
Anthropic holds complex workloads
β†’ Price premium widens (4.4Γ— vs 3.4Γ— in June)
β†’ Revenue share stays at 65%
β†’ No cheap alternative for complex tasks

TREND 3: Open-weight models moving upmarket
────────────────────────────────────────────
Kimi K3 and GLM 5.2 enter agent work
β†’ First open-weight revenue breakthrough
β†’ Open-weight spend share doubles
β†’ Frontier labs' combined share falls below 90%

The Market Segmentation Map

HIGH CAPABILITY / HIGH PRICE
        β”‚
        β”‚  Anthropic (Claude Opus, Fable 5)
        β”‚  β†’ Complex agents, coding
        β”‚
        β”‚  OpenAI (GPT-5 family)
        β”‚  β†’ Mixed workloads
        β”‚
        β”‚  Kimi K3, GLM 5.2  ← NEW ENTRANTS HERE
        β”‚  β†’ Long-horizon agent work
        β”‚
        β”‚  Google, DeepSeek
        β”‚  β†’ Consumer-facing, simple tasks
        β”‚
LOW CAPABILITY / LOW PRICE

Step 10: Key Takeaways and Mental Models

Mental Model 1: The "Volume β‰  Revenue" Rule

A lab can dominate token volume while capturing almost no revenue if its prices are low enough. DeepSeek is the clearest example.

Mental Model 2: The "Task-First" Buying Decision

Enterprise buyers choose the tier (premium vs. cheap) based on what the task requires, then optimize within that tier on price. This is why price competition is "happening in the part of the market with the least money in it."

Mental Model 3: The "Mix Effect" on Average Price

Average price per token is driven more by which models companies choose than by individual model price changes. 81% of July's tokens ran on models that didn't exist 6 months ago.

Mental Model 4: The "Capability Gap as Moat"

Anthropic's pricing power comes not from technical lock-in but from the gap between what cheap models can do and what its customers need. When that gap closes, so does the premium.


Self-Assessment Questions

Test your understanding with these questions:

  1. Basic: If volume grows 59% and spend grows 37%, by approximately how much does price per token change? (Hint: use the formula)

  2. Intermediate: Why did DeepSeek gaining 25% of token volume not significantly hurt Anthropic's 65% revenue share?

  3. Intermediate: What made Kimi K3 different from previous open-weight models like DeepSeek in terms of market impact?

  4. Advanced: The article says "price competition is happening in the part of the market with the least money in it." Explain what this means and why it is true based on the data.

  5. Advanced: If you were advising a startup building an AI-powered coding assistant, which lab would the data suggest you use, and why? What would change your answer if you were building a consumer chatbot instead?


Answer Key

  1. Price per token fell approximately 13.6% β€” volume grew faster than spend, so the average cost per unit dropped.

  2. DeepSeek's tokens are extremely cheap (1/16th of average price), so even massive volume generates little revenue. Anthropic's customers need capabilities that DeepSeek cannot provide, so they don't switch.

  3. Previous open-weight models (like DeepSeek) competed only on price for simple tasks. Kimi K3 entered agent work β€” complex, multi-step tasks β€” at 11Γ— DeepSeek's price, capturing revenue that previously only went to closed-weight labs.

  4. The cheap tier (open-weight, DeepSeek) has most of the volume but almost none of the money. Price wars there don't affect the majority of AI spending, which stays concentrated in Anthropic's premium tier.

  5. Coding assistant: Use Anthropic β€” it collects 80%+ of coding agent spend, suggesting it performs best for this task. Consumer chatbot: Consider DeepSeek or Google β€” high volume, low cost, sufficient for simpler consumer interactions.

More to study