AI Inference Prices Fell 43% Since May. Four Things That Number Hides

AI Inference Prices Fell 43% Since May. Four Things That Number Hides

Key Takeaways

  • The blended enterprise inference price fell 43% in about ten weeks: $2.04 per million tokens on May 31 → $1.45 in late July → $1.16–$1.18 on Aug 6–8, 2026 (Jefferies research, citing Silicon Data’s pricing index, via SCMP).
  • The cuts are uneven. OpenAI reportedly cut GPT-5.6 pricing by up to 80%; Anthropic’s Claude Opus 5 matches its flagship Fable 5 at roughly half the price; DeepSeek’s V4-Flash-0731 runs some tasks at $0.03.
  • Silicon Data’s own index shows frontier models at $4.20 per million tokens and open-weight models at $0.85 — a roughly 5x gap that a single blended average erases.
  • A falling average says nothing about your bill unless you know which side of that gap your workload sits on.

Enterprise AI inference got a lot cheaper on paper this month. The average price of a million tokens dropped from $2.04 in late May to $1.16–$1.18 in the first week of August — a 43% decline in about ten weeks.

That figure comes from Jefferies research built on Silicon Data’s pricing index, reported by the South China Morning Post on August 12. It’s the lowest reading the index has recorded this year.

It’s also a blended average across 400+ models and more than 20 pricing sources. Averages compress. Before assuming your own AI bill is about to shrink by the same 43%, it helps to know what the compression hides — and how that compares with the total-spend pattern we tracked in when AI capex actually pays back.

43% — Enterprise AI token price drop, May 31 to Aug 8, 2026

What Actually Fell, and What Didn’t

Three separate things happened at once this summer, and headlines usually merge them into one number.

WhatPrice pointSource
Blended index (all models)$1.16–$1.18 / M tokens, Aug 6–8Silicon Data via SCMP
Frontier segment$4.20 / M tokens, Apr 22 readingSilicon Data index page
Open-weight segment$0.85 / M tokens, Apr 22 readingSilicon Data index page

We divided Silicon Data’s two segment figures against each other and calculated a frontier-to-open-weight gap of roughly 5x. A company that mostly calls GPT-5.6-class or Claude Opus-class models for accuracy-sensitive work sits near the $4.20 end. A company running high-volume, lower-stakes tasks on open-weight models sits near $0.85.

The $1.16–$1.18 blended figure describes neither company’s actual bill. It describes the market’s shifting mix between the two — the same shift we saw play out in decentralized AI inference pricing, where a cheaper quote didn’t always survive contact with real workloads.

Average enterprise inference price, per million tokens

Individual vendor cuts widen the confusion further. OpenAI’s reported up-to-80% cut on GPT-5.6 pricing doesn’t mean every GPT-5.6 call got 80% cheaper. “Up to” language typically covers the largest-discount tier — often high-volume batch or cached requests — not the median call.

Anthropic’s Claude Opus 5 pricing move is a like-for-like comparison against Fable 5, not against Opus 5’s own prior price. DeepSeek’s $0.03-per-task figure for V4-Flash-0731 is a per-task cost for a specific lightweight model, not a per-million-token rate comparable to the other three. Three different kinds of “cheaper,” reported as one trend.

What Does a Falling Average Price Hide?

  1. The tier you’re actually billed on. A blended average moving down doesn’t move your invoice unless your workload’s tier moved down with it. Check your provider’s price sheet for your specific model, not the index headline.
  2. “Up to” is a ceiling, not a median. Vendor-announced cuts describe their best case. Ask what discount applies to your actual usage pattern — batch, cached, or standard-rate calls each price differently.
  3. Volume can grow faster than price falls. Our earlier reporting on AI capex found that spending keeps climbing even as unit costs drop, because usage volume outpaces the price decline. A 43% price cut paired with a 60% usage increase is still a bigger bill.
  4. New premium tiers launch alongside the cuts. Reasoning and agentic modes are frequently billed as separate, higher-priced tiers even while a provider’s headline model gets cheaper. A falling average can coexist with a rising bill for the features you’ve just started using.

None of this means the price war is fake. Real cuts are happening, and they compound. A team that can shift eligible workloads from frontier to open-weight models captures a real, large saving.

But that saving comes from making a deliberate tier decision, not from the market average moving on its own. The index tells you the market is getting cheaper. It doesn’t tell you whether you are.

A practical version of this checklist takes five minutes. Pull your provider’s invoice for the model you actually call, and compare that line item’s price today against 90 days ago — before comparing it with the index headline.

If your line item didn’t move, the market average moving doesn’t matter to you yet. It’s still worth checking again next quarter, since providers tend to pass frontier-tier cuts through with a lag.

Two Lenses

Why Buying More AI Feels Rational Right Now

Falling average prices are the reason teams keep approving new AI use cases — the math looks easier every quarter. That’s a legitimate read of the data: unit economics for the same task, on the same model tier, genuinely improved through 2026.

Why the Invoice Won’t Match the Headline

The same team that approved three new use cases because “AI got 43% cheaper” can end this quarter with a larger bill than last quarter. Three new use cases at a higher tier cost more than one use case at a lower tier ever did, even after the discount.

Both things are true at once. That’s exactly why the index number and the invoice number diverge, and why finance teams keep asking for a bill that a market-wide average can’t produce.

What would change our view

If Silicon Data or a comparable tracker starts publishing a volume-weighted total-spend figure — not just a per-token price average — and that total falls in step with the per-token price, that would be a real signal the price war is reaching total enterprise AI budgets, not just unit prices.

Until a spend-weighted index exists publicly, a falling average price index tells us about pricing pressure between vendors. It does not tell us what any specific company will pay next quarter.

FAQ

Q. Is AI actually getting cheaper in 2026?

A. The per-token price is falling — down 43% since May 2026 by Jefferies’ measure. Whether your bill falls depends on which model tier you use and how much your usage grows in the same period.

Q. Why doesn’t a lower average token price show up on my invoice?

A. The average blends frontier models (~$4.20/M tokens) and open-weight models (~$0.85/M tokens). If your workload’s mix between the two didn’t shift, the market average moving doesn’t move your specific rate.

Q. Which AI model had the biggest price cut in 2026?

A. OpenAI reportedly cut GPT-5.6 pricing by up to 80%, the largest single-vendor cut reported in this round, though the size of the discount applied to any specific customer’s usage will vary by tier.

Sources

Related from 2mind

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *