Why AI Model Prices Keep Falling and What It Means for Startups
Frontier-model API prices have dropped year over year since 2023. Here's what's actually driving the decline and how startups should plan around it.
In this story 5 sections
AI model prices keep falling because three forces compound at once: labs are shipping smaller models that match older flagship quality, specialized inference chips are getting cheaper per unit of compute, and a growing field of competing labs keeps undercutting each other on price. Together, those forces have pushed the cost per million tokens down by an order of magnitude or more since 2023 for comparable model quality.
If you build a product on top of an API, the sticker price of a model matters as much as the model's capability. At Emergent Wire, we track compute economics the way other outlets track funding rounds, because for a startup, the unit cost of a token is closer to a cost of goods sold than a line item you can ignore. Watching that line fall two or three times in eighteen months changes what's viable to build.
This piece breaks down the three real drivers behind falling AI model prices, what the trend has actually looked like since 2023, and how a startup should think about planning around a cost curve that keeps moving under its feet. It's written for founders and product teams building on top of AI APIs rather than training their own models.
- Model prices for comparable quality have fallen sharply since 2023, driven by efficiency, not just competition.
- Smaller, distilled models now match capability levels that needed a much larger model two years ago.
- Specialized AI chips are cutting the hardware cost side of the equation too.
- A startup's unit economics should be built on current prices, not a bet on future discounts.
- Cheaper isn't automatically better — benchmark against the actual task before switching models.
- Compute costs are still a real constraint even as per-token prices drop, because usage volume keeps rising.
The Price Drop So Far
The scale of the drop is easy to understate if you weren't tracking it month to month. A frontier-quality model that cost several dollars per million output tokens in early 2023 now has a same-tier successor priced at a fraction of that, and a smaller model matching the old flagship's benchmark scores costs less still. Epoch AI, a research organization that tracks compute and cost trends across AI labs, has documented this pattern across multiple model generations: the price for a fixed capability level falls roughly tenfold every one to two years, even as the top-end capability keeps climbing at the same time.
That's an unusual combination. In most industries, the cutting edge stays expensive while only the trailing tier gets cheap. In AI, the trailing tier gets cheap while the cutting edge also gets cheaper for the same job it used to do at a higher price. We've watched this play out directly in our own testing at Emergent Wire: a model we'd have called mid-tier eighteen months ago now runs at prices that would have been considered a rounding error back then.
Efficiency, Not Just Competition
Price wars explain some of the drop, but not most of it. The bigger driver is that labs have gotten much better at squeezing the same output quality out of a smaller, cheaper-to-run model. Distillation, where a smaller model is trained to mimic a larger one's behavior on a narrower set of tasks, is the clearest example. Our explainer on how model distillation works covers the mechanics in more detail, but the short version is that a distilled model can hit 90 percent or more of a larger model's task performance while costing a fraction as much to run.
Open-weight models add a second kind of pressure on price. When a strong open-weight model is available to self-host or run through a low-margin third-party API, closed-model providers have to price competitively against that alternative or lose the price-sensitive part of the market. Our comparison of open-weight versus closed AI models gets into how that competitive dynamic has shifted the market over the past year specifically.
Emergent Wire's read on the compute-economics side is that efficiency gains, not just rivalry between labs, are the more durable driver here. Competition can stall if the market consolidates. Efficiency gains from better training methods and better model architectures tend to stick once they're discovered, which is why the price curve has kept bending down even during periods when only a handful of labs were realistically competing at the frontier.
The Hardware Side of the Curve
The other half of the story is what a token actually costs to compute, separate from what a lab chooses to charge for it. Specialized AI inference chips, both from established chipmakers and from labs building custom silicon, have driven down the raw compute cost per token by shipping more efficient hardware at greater volume. McKinsey & Company, the management consulting firm that tracks data center and compute capital spending, has estimated the scale of investment now going into that hardware buildout, underscoring how much capital sits behind each cheaper token. That's a slower-moving trend than pure software efficiency, but it's a real and additive one.
We track this because it sets a floor under how low prices can realistically go before a provider is running at a loss. Our breakdown of the AI compute buildout and its power demands covers the capital side of this: the data centers and power contracts that have to exist before any of this cheaper inference is even possible at scale, and why that buildout is its own constraint even as per-token prices fall.
| Driver | What It Does to Price | How Durable It Is |
|---|---|---|
| Model distillation | Cuts compute per query for similar quality | High — compounds with each generation |
| Open-weight competition | Forces closed labs to price competitively | Medium — depends on open-weight quality staying close |
| Specialized inference chips | Lowers raw compute cost per token | High but slower-moving |
| Lab price competition | Compresses margins directly | Lower — can stall with market consolidation |
What It Means for Startups
For a startup, falling model prices are a genuine tailwind, but they're an unreliable one to build a business plan around directly. We've seen founders price a product assuming next year's token cost rather than this year's, and get burned when a price cut arrives later than expected or a specific model they depend on gets deprecated instead of discounted.
A more durable approach: build unit economics on today's prices, and treat any drop as margin upside rather than a baked-in assumption. Where possible, keep the product's logic portable across models rather than hard-wired to one provider's pricing tier, since the cheapest reasonable option for a given task keeps changing every few months. A startup that architected itself around one specific model's cost structure a year ago is often now stuck migrating rather than benefiting from the very price drops that should have helped it.
Task-appropriate model selection matters more than chasing the single cheapest option available. A smaller, cheaper model that gets a task wrong 15 percent more often than a pricier one can end up costing more once you account for the manual review or customer complaints that follow. At Emergent Wire, we'd rather see a team benchmark two or three model tiers against their actual production data than assume price alone tells the whole story.
The Bottom Line
Model prices are falling because efficiency, hardware, and competition are all pulling in the same direction at once, and there's no strong reason to expect that to reverse soon. For a startup, the right response isn't to bet the business plan on next year's discount — it's to build with today's costs, stay flexible about which model does which job, and treat every price drop as it arrives rather than in advance.
Emergent Wire covers the AI industry's economics, from compute costs to the deals shaping who can afford to build what, for readers who want the numbers behind the headlines.