On Friday 24 July, Anthropic released Claude Opus 5 and priced it at $5 per million input tokens and $25 per million output — in the company's own words, close to the intelligence of its flagship Fable 5 model at half the price. Three days earlier, Google had cut the sticker on its workhorse Gemini Flash line roughly in half. Two days after that, China's Moonshot AI uploaded the largest open-weight model ever built for anyone to download. If you priced an AI project for your business a year ago and shelved it because the numbers didn't work, that quote is now out of date.
One week, three price signals
Anthropic. Claude Opus 5 landed on 24 July at the same rate as the Opus 4.8 it replaces, but with markedly more capability — Axios reports it approaches Anthropic's top-end Fable 5 on many tasks, at half that model's price. It is the fourth Claude 5 model in under two months, it becomes the default for Claude Max subscribers, and it ships with an "effort" dial that trades a little quality for a lot less spend on routine work.
Google. On 21 July, Gemini 3.6 Flash arrived with introductory pricing of $0.75 per million input tokens and $3.75 per million output — a fraction of the $1.50 and $9.00 that Gemini 3.5 Flash charged at its launch in May. The fine print matters: Google says the introductory rate expires on 31 December 2026, after which it doubles.
Moonshot AI. Days after the Opus launch, Moonshot published the weights of Kimi K3 on Hugging Face: 2.8 trillion parameters, a one-million-token context window, and a modified MIT licence that permits free commercial use for almost any normal business. Coverage widely called it the biggest open-weight model in history. You will probably never run it yourself — self-hosting requires roughly 1.5 terabytes of storage — but its existence keeps pressure on everyone else's price list.
Add OpenAI's GPT-5.6, released on 9 July in three tiers, and the pattern is clear. This is not one vendor running a promotion. It is the whole market repricing, again.
The curve behind the headlines
Stanford's AI Index tracked what it costs to buy a fixed level of capability — performance equivalent to the original ChatGPT — and found the price fell from $20 per million tokens in November 2022 to $0.07 by October 2024. That is a 280-fold drop in under two years. The same report notes that, depending on the task, LLM inference prices have been falling anywhere from 9 to 900 times per year, and the research group Epoch AI puts the decline at roughly 40 times per year for GPT-4-level quality. Competition does part of the work: China's DeepSeek offers its R1 reasoning model at $0.55 in and $2.19 out per million tokens, a small fraction of what Western reasoning models charged a year earlier.
The direction is one-way, but individual price tags are not. Google's expiring introductory rate is a reminder that "cheap" sometimes means "cheap until January". The honest summary: any given capability gets dramatically cheaper over time, while vendors keep finding new, more expensive things to sell at the top.
What now pencils out
A million tokens is roughly 750,000 words. At mid-2026 prices, some projects that looked extravagant in 2025 become rounding errors.
Customer email, drafted for you. Say your webshop in Ghent answers 2,000 customer messages a month, and a model reads about 600 tokens and writes about 400 per reply. That is 1.2 million input and 0.8 million output tokens: about $26 a month on a frontier model like Opus 5, and under $4 on Gemini 3.6 Flash. The bottleneck is no longer the model bill — it is connecting the model to your inbox and deciding what it may send.
Product texts in three languages. Rewriting 500 product descriptions in French, German and Dutch is on the order of 600,000 output tokens — a few dollars on a budget model, perhaps $15 to $20 on a frontier one. A year ago the same job on frontier quality was a line item; now it is a coffee run.
Meetings and calls. Summarising every sales call and site visit — hours of transcript per week — costs cents at Flash prices. Tasks with high volume and low stakes are exactly where the price collapse bites first.
Subscription or API?
For a small business the split is simple. People use subscriptions. A flat plan in the €20-a-month range gives you and your staff a chat interface, file handling and the newest models, with a predictable bill — for daily interactive use it is nearly always the better deal. Software uses the API. The moment a task runs without a human in the loop — drafting replies, tagging orders, summarising calls — pay per token. As the arithmetic above shows, typical SMB volumes now cost single-digit euros a month there. Many businesses sensibly run both: one subscription for the owner, one API key behind one automated workflow.
Don't marry a vendor
Four Claude 5 models in under two months, per Axios. Two Google price moves in one summer. A record open model out of Beijing between them. In a market that leapfrogs every few weeks, the expensive mistake is wiring your business to a single supplier's model and price list. Keep your prompts and instructions documented somewhere portable, prefer tools that let you switch the underlying model, and re-quote your workloads every quarter the way you would an energy contract. And never build a business case on an introductory rate with a published expiry date.
When cheaper doesn't help
Falling token prices fix exactly one line of a project budget. Integration work, cleaning up your data, and the hours your team spends changing how it works usually cost more than the model ever will, and none of that got cheaper in July. If a task has no volume — a handful of quotes a week — the savings are pennies, and the right tool is the €20 subscription you may already have. If you piloted something last year and quality was the problem, a price cut solves nothing; retest capability instead. And your GDPR homework — what data leaves your systems, under which agreement — is identical at every price point.
What the July price wave really changes is which ideas are worth testing: the shortlist you wrote a year ago was priced against a market that no longer exists. If you want to figure out where AI genuinely fits in your business before spending anything at all, Cresly's AI Readiness Scan maps your processes against what current models can do — at prices that, as of this month, are half of what you probably assumed.