AI inference costs are falling fast. Here is why ChatGPT, Claude, Gemini, and AI tools can still have pricey plans and strict limits.
AI Is Getting Cheaper Fast. Why Are Your Limits Still There?If AI is getting cheaper so quickly, why does a useful subscription still come with a price tag, a rate limit, and the occasional “try again later” message?
It is a fair question. The headlines point in one direction: smaller models are getting better, hardware is getting more efficient, and API prices keep finding new ways to be lower than last year’s. The experience of a subscriber often points in another: premium plans, message caps, slower access at busy times, and features that consume a limit surprisingly quickly.
Both observations can be true because the cost of a token is not the cost of a product.
The useful way to read the cost curve is not “AI should now be free.” It is “more useful intelligence is becoming cheap enough to put inside more workflows.” That is a much bigger change—and it comes with its own engineering and business constraints.
The best clean comparison I have found comes from Stanford’s 2025 AI Index. It estimates that the cost to query a model with GPT-3.5-level MMLU performance fell from $20 per million tokens in November 2022 to $0.07 per million tokens in October 2024. That is more than a 280× reduction.
This is not a claim that every model, task, or subscription became 280 times cheaper. It compares systems at a roughly comparable capability point on one benchmark. That distinction matters. Still, it is a serious shift in what can be offered economically.

Two reference points, not a subscription-price series:
The chart uses a logarithmic axis because the two values are so far apart. On an ordinary linear chart, the lower-cost bar would be nearly invisible. The point is not a gentle discount; it is a very large drop at a comparable capability reference point.
\nThe same report attributes the change to a mix of better hardware, better models, and more efficient inference. It says hardware costs declined about 30% annually and energy efficiency improved about 40% annually in the period it studied. It also notes a useful size comparison: the smallest model above 60% on MMLU went from PaLM at 540 billion parameters in 2022 to Phi-3-mini at 3.8 billion in 2024.
That is the important trend: capability is not only coming from making one giant model larger. It is also coming from better architectures, data, post-training, distillation, quantization, routing, and systems work.
“AI cost” bundles several different things together. Separating them prevents most bad arguments about pricing.
For an API user, a token price is the visible unit. It is genuinely useful: OpenAI’s pricing page separates input, cached input, and output pricing, while Google documents lower-cost Batch, Flex, and caching paths for workloads that can tolerate different latency and reliability trade-offs.
But the final bill depends on the shape of the work. A short prompt that starts a long reasoning process, calls several tools, searches files, creates an image, or loops through an agent workflow is not comparable to one quick text completion. A long context can be inexpensive when it is reused from a cache, or expensive when every turn changes it enough to miss the cache.
Caching is a good example. If a product repeatedly sends the same large system prompt, document, or codebase context, it can reuse that work. OpenAI describes discounted cached input on its pricing page; Google lists caching alongside explicit storage and latency trade-offs. Batch processing can cut costs further when an answer does not need to arrive immediately. Google’s current optimization guide, for example, lists Batch at 50% of standard pricing and a 90% cache-read discount, subject to the product’s specific terms.
Those are not magic coupons. They work best when the request has a reusable prefix or can wait in a queue. A live conversation that changes context every turn, a real-time coding agent, or an interactive image request is a different workload.
That is why good AI systems increasingly route work rather than send every request to the biggest available model. A small model may classify an email, extract a field, or summarize a stable template. A stronger model may handle the ambiguous step. The product gets cheaper without pretending every decision deserves the same amount of compute.
The practical question for a developer is not “Which model is cheapest?” It is “Which part of this workflow needs which level of quality, latency, context, and verification?” That is the same judgment boundary I described in Use AI to Become a Better Software Developer, Not Just a Faster One: speed is useful only when the decision-making around it stays honest.
A subscription is not a metered token invoice with a nicer logo.
It typically offers access to a collection of workloads with wildly different cost profiles: short chat answers, long-context analysis, extended reasoning, file processing, browsing, image generation, tool calls, coding agents, synchronization, storage, and support. The expensive part is often not the average request. It is the tail: a customer who asks for a deep answer over a huge context at the same busy moment as everyone else.
Three constraints follow from that.
An interactive product has to be responsive when its customers arrive, not only when machines happen to be idle. Providers therefore have to reserve or obtain capacity for busy periods, protect service quality, and prevent a small number of intensive sessions from making the product unusable for everyone else.
That is what a rate limit is often doing: managing a shared, time-sensitive resource. It is not evidence that inference prices never fell.
Reasoning models make this easy to miss. An answer may look short while the system used substantial hidden work to plan, check, or search. Some providers expose effort controls precisely because higher effort trades latency and cost for quality. Anthropic’s Claude 3.7 announcement is explicit about letting API users set a thinking-token budget for that trade-off.
The same applies to agentic work. “Fix this bug” may involve reading files, searching documentation, calling tools, running tests, retrying failed steps, and holding a large context. The unit that matters is no longer a single prompt; it is a workflow.
This is the part that makes the change feel paradoxical.
When a task becomes cheap enough, people do not merely do the old task at a lower bill. They give the model more context, ask for more revisions, add verification, automate a recurring process, switch from text to images or video, and let an agent stay on the task longer.
That is normal economic behavior. The cost per unit falls; the number and ambition of units rises.
For a developer, the equivalent is moving from “explain this stack trace” to “inspect the live browser, trace the network call, propose a patch, run targeted tests, and summarize the evidence.” That workflow can be far more valuable, but it is also not a one-token question. Let Your AI Agent Debug the Browser Instead of Guessing makes the same point from the engineering side: tools produce evidence, but every tool call and verification step is real work.
The product-level consequence is straightforward: a lower unit cost lets a product put more intelligence inside one user action. That action may now include:
That is why a plan can still need limits even while the underlying unit cost falls.
This should make us cautious in both directions.
It is wrong to look at a cheap small-model API price and conclude that a general-purpose AI product should be free. It is also wrong to look at a subscription cap and conclude that the underlying technology is not getting cheaper. They describe different layers of the stack.
No. A falling inference-cost curve cannot tell you whether any particular company, stock, or funding round is priced sensibly. Those are valuation questions, and they require revenue, margins, competition, concentration risk, capital expenditure, and a lot more than a token chart.
What the curve does tell us is narrower and more practical: a capability that was too expensive for a routine workflow can eventually cross a threshold where it is worth embedding.
Speech-to-text is a familiar example of this pattern. Once it became good and cheap enough, it moved from a specialist service into keyboards, meeting tools, accessibility features, and operating systems. The interesting question was not whether every company in the space deserved the same valuation. It was which workflows changed once the unit economics became acceptable.
AI is doing something similar, but with more ambiguity and more risk. Cheap inference does not remove hallucinations, weak requirements, privacy constraints, copyright questions, vendor dependency, or the need for human verification. It lowers the cost of attempting useful work. It does not lower the cost of being wrong to zero.
You do not need to turn every team into an AI startup. You also do not need to hand important decisions to a chatbot because a demo looked impressive.
But it is worth learning where the economics have already changed your work:
The teams that benefit will not be the ones with the longest prompt libraries. They will be the ones that can direct a model, constrain a workflow, evaluate evidence, and choose when not to automate.
That is a learnable skill. And as the cost of useful intelligence keeps falling, it will stop being a niche skill and become part of the baseline expectation for many knowledge-work roles.
Before you add AI to a task, write down four things:
That turns the conversation from “Why is this plan expensive?” into a better engineering question: “Which capability are we buying, what workload creates the cost, and how can we use it deliberately?”
AI is getting cheaper. The right response is neither blind hype nor dismissal. It is to get better at turning a falling unit cost into a workflow that is genuinely useful—and still worth trusting.
Where in your work would a much cheaper model change the workflow itself, rather than simply make the old workflow faster?