Analysis

Why AI Model Prices Keep Falling and What It Means for Startups

Frontier-model API prices have dropped year over year since 2023. Here's what's actually driving the decline and how startups should plan around it.

Tobias Reyes

AI Industry & Policy Analyst

Published 7 min read
Yellow letter tiles spell the word 'price' against a vibrant blue backdrop, ideal for business concepts.
In this story 5 sections

AI model prices keep falling because three forces compound at once: labs are shipping smaller models that match older flagship quality, specialized inference chips are getting cheaper per unit of compute, and a growing field of competing labs keeps undercutting each other on price. Together, those forces have pushed the cost per million tokens down by an order of magnitude or more since 2023 for comparable model quality.

If you build a product on top of an API, the sticker price of a model matters as much as the model's capability. At Emergent Wire, we track compute economics the way other outlets track funding rounds, because for a startup, the unit cost of a token is closer to a cost of goods sold than a line item you can ignore. Watching that line fall two or three times in eighteen months changes what's viable to build.

This piece breaks down the three real drivers behind falling AI model prices, what the trend has actually looked like since 2023, and how a startup should think about planning around a cost curve that keeps moving under its feet. It's written for founders and product teams building on top of AI APIs rather than training their own models.

  • Model prices for comparable quality have fallen sharply since 2023, driven by efficiency, not just competition.
  • Smaller, distilled models now match capability levels that needed a much larger model two years ago.
  • Specialized AI chips are cutting the hardware cost side of the equation too.
  • A startup's unit economics should be built on current prices, not a bet on future discounts.
  • Cheaper isn't automatically better — benchmark against the actual task before switching models.
  • Compute costs are still a real constraint even as per-token prices drop, because usage volume keeps rising.

The Price Drop So Far

The scale of the drop is easy to understate if you weren't tracking it month to month. A frontier-quality model that cost several dollars per million output tokens in early 2023 now has a same-tier successor priced at a fraction of that, and a smaller model matching the old flagship's benchmark scores costs less still. Epoch AI, a research organization that tracks compute and cost trends across AI labs, has documented this pattern across multiple model generations: the price for a fixed capability level falls roughly tenfold every one to two years, even as the top-end capability keeps climbing at the same time.

That's an unusual combination. In most industries, the cutting edge stays expensive while only the trailing tier gets cheap. In AI, the trailing tier gets cheap while the cutting edge also gets cheaper for the same job it used to do at a higher price. We've watched this play out directly in our own testing at Emergent Wire: a model we'd have called mid-tier eighteen months ago now runs at prices that would have been considered a rounding error back then.

Stock market data chart showing trends in red and green. Perfect for financial and business themes.

Efficiency, Not Just Competition

Price wars explain some of the drop, but not most of it. The bigger driver is that labs have gotten much better at squeezing the same output quality out of a smaller, cheaper-to-run model. Distillation, where a smaller model is trained to mimic a larger one's behavior on a narrower set of tasks, is the clearest example. Our explainer on how model distillation works covers the mechanics in more detail, but the short version is that a distilled model can hit 90 percent or more of a larger model's task performance while costing a fraction as much to run.

Open-weight models add a second kind of pressure on price. When a strong open-weight model is available to self-host or run through a low-margin third-party API, closed-model providers have to price competitively against that alternative or lose the price-sensitive part of the market. Our comparison of open-weight versus closed AI models gets into how that competitive dynamic has shifted the market over the past year specifically.

Emergent Wire's read on the compute-economics side is that efficiency gains, not just rivalry between labs, are the more durable driver here. Competition can stall if the market consolidates. Efficiency gains from better training methods and better model architectures tend to stick once they're discovered, which is why the price curve has kept bending down even during periods when only a handful of labs were realistically competing at the frontier.

Detailed macro view of a circuit board showcasing microchips and electronic components.

The Hardware Side of the Curve

The other half of the story is what a token actually costs to compute, separate from what a lab chooses to charge for it. Specialized AI inference chips, both from established chipmakers and from labs building custom silicon, have driven down the raw compute cost per token by shipping more efficient hardware at greater volume. McKinsey & Company, the management consulting firm that tracks data center and compute capital spending, has estimated the scale of investment now going into that hardware buildout, underscoring how much capital sits behind each cheaper token. That's a slower-moving trend than pure software efficiency, but it's a real and additive one.

We track this because it sets a floor under how low prices can realistically go before a provider is running at a loss. Our breakdown of the AI compute buildout and its power demands covers the capital side of this: the data centers and power contracts that have to exist before any of this cheaper inference is even possible at scale, and why that buildout is its own constraint even as per-token prices fall.

The Hardware Side of the Curve
DriverWhat It Does to PriceHow Durable It Is
Model distillationCuts compute per query for similar qualityHigh — compounds with each generation
Open-weight competitionForces closed labs to price competitivelyMedium — depends on open-weight quality staying close
Specialized inference chipsLowers raw compute cost per tokenHigh but slower-moving
Lab price competitionCompresses margins directlyLower — can stall with market consolidation
Three colleagues working together on a laptop, fostering teamwork and innovation.

What It Means for Startups

For a startup, falling model prices are a genuine tailwind, but they're an unreliable one to build a business plan around directly. We've seen founders price a product assuming next year's token cost rather than this year's, and get burned when a price cut arrives later than expected or a specific model they depend on gets deprecated instead of discounted.

A more durable approach: build unit economics on today's prices, and treat any drop as margin upside rather than a baked-in assumption. Where possible, keep the product's logic portable across models rather than hard-wired to one provider's pricing tier, since the cheapest reasonable option for a given task keeps changing every few months. A startup that architected itself around one specific model's cost structure a year ago is often now stuck migrating rather than benefiting from the very price drops that should have helped it.

Task-appropriate model selection matters more than chasing the single cheapest option available. A smaller, cheaper model that gets a task wrong 15 percent more often than a pricier one can end up costing more once you account for the manual review or customer complaints that follow. At Emergent Wire, we'd rather see a team benchmark two or three model tiers against their actual production data than assume price alone tells the whole story.

The Bottom Line

Model prices are falling because efficiency, hardware, and competition are all pulling in the same direction at once, and there's no strong reason to expect that to reverse soon. For a startup, the right response isn't to bet the business plan on next year's discount — it's to build with today's costs, stay flexible about which model does which job, and treat every price drop as it arrives rather than in advance.

Emergent Wire covers the AI industry's economics, from compute costs to the deals shaping who can afford to build what, for readers who want the numbers behind the headlines.

Why has the price of AI model tokens dropped so much?
Three things stack together: model efficiency gains that cut the compute needed for a given quality level, cheaper specialized chips shipping in volume, and direct price competition between a growing number of frontier labs. Each factor alone would matter; together they compound year over year.
Will AI model prices keep falling at the same rate?
Probably not at the exact same rate forever, but the trend is likely to continue for now. Hardware costs and efficiency techniques both still have real room to improve, and competitive pressure between labs hasn't let up. A slowdown is more likely than a reversal in the near term.
Should a startup build on the cheapest available model?
Not automatically. Cheaper models are usually fine for high-volume, low-complexity tasks like classification or summarization, but a startup should still benchmark the cheaper option against its actual task before switching, since a lower price with a higher error rate can cost more in cleanup than it saves.
How should a startup plan its budget around falling AI prices?
Build unit economics around today's prices, not a hoped-for future discount, and treat any price drop as a margin improvement rather than a reason to scale usage indiscriminately. Locking a product's pricing to a specific model's cost structure is risky given how fast that structure keeps shifting.