A downward orange cost curve passes over progressively smaller computer chips and stacks of coins on a light technical workspace.
Mohamad Abuzaid 8 hours ago
mohamad-abuzaid #ai

AI Is Getting Cheaper Fast. Why Are Your Limits Still There?

AI inference costs are falling fast. Here is why ChatGPT, Claude, Gemini, and AI tools can still have pricey plans and strict limits.

AI Is Getting Cheaper Fast. Why Are Your Limits Still There?

If AI is getting cheaper so quickly, why does a useful subscription still come with a price tag, a rate limit, and the occasional “try again later” message?

It is a fair question. The headlines point in one direction: smaller models are getting better, hardware is getting more efficient, and API prices keep finding new ways to be lower than last year’s. The experience of a subscriber often points in another: premium plans, message caps, slower access at busy times, and features that consume a limit surprisingly quickly.

Both observations can be true because the cost of a token is not the cost of a product.

The useful way to read the cost curve is not “AI should now be free.” It is “more useful intelligence is becoming cheap enough to put inside more workflows.” That is a much bigger change—and it comes with its own engineering and business constraints.

The headline is real: comparable inference got dramatically cheaper

The best clean comparison I have found comes from Stanford’s 2025 AI Index. It estimates that the cost to query a model with GPT-3.5-level MMLU performance fell from $20 per million tokens in November 2022 to $0.07 per million tokens in October 2024. That is more than a 280× reduction.

This is not a claim that every model, task, or subscription became 280 times cheaper. It compares systems at a roughly comparable capability point on one benchmark. That distinction matters. Still, it is a serious shift in what can be offered economically.

Logarithmic horizontal bar chart showing comparable AI inference cost falling from 20 dollars per million tokens in November 2022 to 0.07 dollars in October 2024.

Two reference points, not a subscription-price series:

  • November 2022: GPT-3.5-level reference — $20.00 per million tokens.
  • October 2024: Gemini 1.5 Flash-8B at the same MMLU capability reference point — $0.07 per million tokens.

The chart uses a logarithmic axis because the two values are so far apart. On an ordinary linear chart, the lower-cost bar would be nearly invisible. The point is not a gentle discount; it is a very large drop at a comparable capability reference point.

\n

The same report attributes the change to a mix of better hardware, better models, and more efficient inference. It says hardware costs declined about 30% annually and energy efficiency improved about 40% annually in the period it studied. It also notes a useful size comparison: the smallest model above 60% on MMLU went from PaLM at 540 billion parameters in 2022 to Phi-3-mini at 3.8 billion in 2024.

That is the important trend: capability is not only coming from making one giant model larger. It is also coming from better architectures, data, post-training, distillation, quantization, routing, and systems work.

What exactly got cheaper?

“AI cost” bundles several different things together. Separating them prevents most bad arguments about pricing.

  • Model inference: The compute required to process input and generate output at a given capability level is becoming cheaper. That does not mean every request uses a small or inexpensive model.
  • API list price: Providers can publish lower prices per input or output token for particular models. That does not create a flat cost per task; output, reasoning, tools, and media change the bill.
  • Hardware efficiency: Accelerators can perform more work per watt and per generation. That does not create infinite capacity at the exact moment everyone wants an answer.
  • Capability per dollar: Smaller or specialized models can reach useful quality. That does not make them equally reliable on every task.
  • Systems efficiency: Caching, batching, routing, and better serving software reduce waste. That does not guarantee savings when work must be interactive, novel, or served at peak time.

For an API user, a token price is the visible unit. It is genuinely useful: OpenAI’s pricing page separates input, cached input, and output pricing, while Google documents lower-cost Batch, Flex, and caching paths for workloads that can tolerate different latency and reliability trade-offs.

But the final bill depends on the shape of the work. A short prompt that starts a long reasoning process, calls several tools, searches files, creates an image, or loops through an agent workflow is not comparable to one quick text completion. A long context can be inexpensive when it is reused from a cache, or expensive when every turn changes it enough to miss the cache.

The efficiency tools are real, but they have conditions

Caching is a good example. If a product repeatedly sends the same large system prompt, document, or codebase context, it can reuse that work. OpenAI describes discounted cached input on its pricing page; Google lists caching alongside explicit storage and latency trade-offs. Batch processing can cut costs further when an answer does not need to arrive immediately. Google’s current optimization guide, for example, lists Batch at 50% of standard pricing and a 90% cache-read discount, subject to the product’s specific terms.

Those are not magic coupons. They work best when the request has a reusable prefix or can wait in a queue. A live conversation that changes context every turn, a real-time coding agent, or an interactive image request is a different workload.

That is why good AI systems increasingly route work rather than send every request to the biggest available model. A small model may classify an email, extract a field, or summarize a stable template. A stronger model may handle the ambiguous step. The product gets cheaper without pretending every decision deserves the same amount of compute.

The practical question for a developer is not “Which model is cheapest?” It is “Which part of this workflow needs which level of quality, latency, context, and verification?” That is the same judgment boundary I described in Use AI to Become a Better Software Developer, Not Just a Faster One: speed is useful only when the decision-making around it stays honest.

Why subscribers do not always feel the drop

A subscription is not a metered token invoice with a nicer logo.

It typically offers access to a collection of workloads with wildly different cost profiles: short chat answers, long-context analysis, extended reasoning, file processing, browsing, image generation, tool calls, coding agents, synchronization, storage, and support. The expensive part is often not the average request. It is the tail: a customer who asks for a deep answer over a huge context at the same busy moment as everyone else.

Three constraints follow from that.

1. Peak capacity costs more than average capacity

An interactive product has to be responsive when its customers arrive, not only when machines happen to be idle. Providers therefore have to reserve or obtain capacity for busy periods, protect service quality, and prevent a small number of intensive sessions from making the product unusable for everyone else.

That is what a rate limit is often doing: managing a shared, time-sensitive resource. It is not evidence that inference prices never fell.

2. Better answers can spend more compute

Reasoning models make this easy to miss. An answer may look short while the system used substantial hidden work to plan, check, or search. Some providers expose effort controls precisely because higher effort trades latency and cost for quality. Anthropic’s Claude 3.7 announcement is explicit about letting API users set a thinking-token budget for that trade-off.

The same applies to agentic work. “Fix this bug” may involve reading files, searching documentation, calling tools, running tests, retrying failed steps, and holding a large context. The unit that matters is no longer a single prompt; it is a workflow.

3. Cheaper work creates more work

This is the part that makes the change feel paradoxical.

When a task becomes cheap enough, people do not merely do the old task at a lower bill. They give the model more context, ask for more revisions, add verification, automate a recurring process, switch from text to images or video, and let an agent stay on the task longer.

That is normal economic behavior. The cost per unit falls; the number and ambition of units rises.

For a developer, the equivalent is moving from “explain this stack trace” to “inspect the live browser, trace the network call, propose a patch, run targeted tests, and summarize the evidence.” That workflow can be far more valuable, but it is also not a one-token question. Let Your AI Agent Debug the Browser Instead of Guessing makes the same point from the engineering side: tools produce evidence, but every tool call and verification step is real work.

A cheap token is not a cheap product

The product-level consequence is straightforward: a lower unit cost lets a product put more intelligence inside one user action. That action may now include:

  • longer context;
  • deeper reasoning;
  • files, images, audio, or video;
  • tool calls and agent loops; and
  • peak-time capacity, reliability, safety, and support work.

That is why a plan can still need limits even while the underlying unit cost falls.

This should make us cautious in both directions.

It is wrong to look at a cheap small-model API price and conclude that a general-purpose AI product should be free. It is also wrong to look at a subscription cap and conclude that the underlying technology is not getting cheaper. They describe different layers of the stack.

Does cheaper AI settle the “bubble” debate?

No. A falling inference-cost curve cannot tell you whether any particular company, stock, or funding round is priced sensibly. Those are valuation questions, and they require revenue, margins, competition, concentration risk, capital expenditure, and a lot more than a token chart.

What the curve does tell us is narrower and more practical: a capability that was too expensive for a routine workflow can eventually cross a threshold where it is worth embedding.

Speech-to-text is a familiar example of this pattern. Once it became good and cheap enough, it moved from a specialist service into keyboards, meeting tools, accessibility features, and operating systems. The interesting question was not whether every company in the space deserved the same valuation. It was which workflows changed once the unit economics became acceptable.

AI is doing something similar, but with more ambiguity and more risk. Cheap inference does not remove hallucinations, weak requirements, privacy constraints, copyright questions, vendor dependency, or the need for human verification. It lowers the cost of attempting useful work. It does not lower the cost of being wrong to zero.

The risky position is treating the curve as someone else’s problem

You do not need to turn every team into an AI startup. You also do not need to hand important decisions to a chatbot because a demo looked impressive.

But it is worth learning where the economics have already changed your work:

  • Which repetitive analysis can be drafted and then checked by a person?
  • Which support, documentation, or internal-search tasks benefit from a smaller, bounded model?
  • Which workflows need a stronger model—and which only need a better prompt, retrieval boundary, or cache strategy?
  • Which decisions must remain explicitly owned by a human because the risk is legal, financial, security-sensitive, or irreversible?

The teams that benefit will not be the ones with the longest prompt libraries. They will be the ones that can direct a model, constrain a workflow, evaluate evidence, and choose when not to automate.

That is a learnable skill. And as the cost of useful intelligence keeps falling, it will stop being a niche skill and become part of the baseline expectation for many knowledge-work roles.

A practical rule for the next AI workflow

Before you add AI to a task, write down four things:

  1. The unit of work: one classification, one draft, one test investigation, one customer reply, or one agent run.
  2. The cost drivers: model tier, context size, output length, tools, media, retries, and required latency.
  3. The quality boundary: what can be automated, what needs review, and what must never be delegated.
  4. The evidence: what proves the workflow helped instead of merely producing more output.

That turns the conversation from “Why is this plan expensive?” into a better engineering question: “Which capability are we buying, what workload creates the cost, and how can we use it deliberately?”

AI is getting cheaper. The right response is neither blind hype nor dismissal. It is to get better at turning a falling unit cost into a workflow that is genuinely useful—and still worth trusting.

Where in your work would a much cheaper model change the workflow itself, rather than simply make the old workflow faster?

Sources

Effective UI/UX Design in Android Apps (3/3)

Effective UI/UX Design in Android Apps (3/3)

1675112374.jpg
Mohamad Abuzaid
2 years ago
Jetpack Compose Animation

Jetpack Compose Animation

1675112374.jpg
Mohamad Abuzaid
2 years ago
Kotlin's Interoperability with Java (2/3)

Kotlin's Interoperability with Java (2/3)

1675112374.jpg
Mohamad Abuzaid
2 years ago
Security in Android App Development (2/3)

Security in Android App Development (2/3)

1675112374.jpg
Mohamad Abuzaid
2 years ago
SOLID principles explained

SOLID principles explained

1675112374.jpg
Mohamad Abuzaid
3 years ago