The AI Industry

AI Inference Costs: What a Million Tokens Actually Costs in 2026

A unit-economics breakdown of what running a million tokens through today's frontier and budget models really costs, and why the sticker price is only half the bill.

Tobias Reyes

AI Industry & Policy Analyst

Published 6 min read
Close-up of Scrabble tiles spelling 'Token' on a wooden surface with a blurred green background.
In this story 5 sections

Quick answer: In September 2026, a million tokens costs anywhere from a few cents to more than $50, depending on model tier and whether you're counting input or output tokens. Budget models like GPT-5 nano charge under $1 per million tokens combined, while frontier reasoning models charge $20 to $30 per million output tokens alone.

Every team running AI in production eventually asks the same question: what does this actually cost per request? The answer used to be simple when there were three or four models to choose from. Now there are dozens of tiers across OpenAI, Anthropic, and Google, each priced differently for input versus output, cached versus fresh context, and standard versus batch processing. This guide walks through real 2026 per-token pricing, why it keeps dropping even as usage bills climb, and how to actually budget for a million tokens instead of guessing. It's written for engineering leads, finance teams, and anyone trying to forecast an AI line item that used to not exist.

Detailed view of a server rack with a focus on technology and data storage.

Real 2026 Pricing: Comparing Model Tiers

At Emergent Wire, we track AI infrastructure economics as one of our core beats, and the spread across today's model tiers is wider than at any point since token-based pricing became standard. The table below reflects standard-tier API rates as of September 2026, not promotional or negotiated enterprise pricing.

Real 2026 Pricing: Comparing Model Tiers
Model tierInput ($/1M tokens)Output ($/1M tokens)Typical use case
GPT-5 nano$0.05$0.40High-volume classification, routing
Claude Haiku 4.5$1.00$5.00Short summaries, extraction
GPT-5 mini$0.25$2.00Lightweight assistants
Claude Sonnet 5$2.00$10.00General-purpose agentic work
GPT-5.6-terra$2.00$12.00Mid-tier reasoning tasks
Claude Opus 5.5$4.00$20.00Long-running agentic coding
GPT-5.6-sol$5.00$30.00Frontier reasoning, complex analysis

The pattern holds across every provider: output tokens cost far more than input tokens, because generating text is more compute-intensive than reading it. A team that shifts even part of a workload from a frontier reasoning model to a mid-tier model like Claude Sonnet 5 can cut its output bill by half or more without necessarily losing much quality for the task at hand.

Detailed macro shot of a CPU microchip with focus on golden pins, highlighting technology details.

Why Inference Prices Keep Falling So Fast

Per-token prices have collapsed since generative AI went mainstream. Epoch AI, a nonprofit research group that tracks AI compute trends, found that the cost of running a model at GPT-3.5's benchmark level fell from about $20 per million tokens in late 2022 to roughly $0.07 per million tokens by late 2024. That's a drop of more than 280-fold in two years, and Epoch AI's broader dataset puts the median annual decline for equivalent capability at around 50x per year since 2022, rising toward 200x per year when measured only from 2024 onward.

Three forces drive that curve. First, architecture improvements like mixture-of-experts designs let a model activate only a fraction of its parameters per token, cutting compute per request without shrinking total model capacity. Second, GPU utilization has improved through better batching and scheduling, so providers serve more requests per chip per hour. Third, competition between OpenAI, Anthropic, and Google has turned pricing into a genuine lever, not just a cost pass-through. A fourth factor gets less attention but matters just as much: newer, smaller models trained on better data now match older flagship models on many benchmarks, at a fraction of the parameter count and inference cost. That "smaller model, same capability" pattern is a big part of why the equivalent-capability price curve keeps bending downward even between major architecture shifts.

We've seen this play out directly in our own testing at Emergent Wire: a task that cost roughly $4 to run through a frontier model in early 2024 now runs for well under a dollar on a comparably capable mid-tier model. That's the Epoch AI curve showing up in a real invoice, not just a research chart. The catch is that this comparison only holds for equivalent capability, not equivalent model tier. A frontier-tier model in 2026 still costs frontier-tier prices; what got cheaper is the price of reaching a fixed capability bar, since that bar keeps getting hit by smaller, more efficient models over time.

A detailed view of a blue lit computer server rack in a data center showcasing technology and hardware.

The Hidden Costs Behind the Sticker Price

Falling per-token prices don't mean falling bills. Enterprise AI spend is climbing fast even as unit costs drop, because the number of tokens consumed per task keeps growing.

Three factors inflate real usage well beyond the sticker price. Reasoning models generate internal "thinking" tokens before producing a visible answer, and those tokens bill at the output rate even though the user never sees them directly. Longer context windows mean every request re-processes more prior conversation or document text as input, which adds up fast on long sessions. And multi-step agent workflows chain several model calls together for a single user-facing task, multiplying token usage several times over compared to a single-turn chat response.

Electricity is a smaller but real line item behind the scenes. According to the U.S. Energy Information Administration (EIA), industrial electricity prices averaged roughly 8.5 to 8.8 cents per kilowatt-hour nationally through mid-2026, though very large data centers frequently negotiate separate utility tariffs that don't match the published average. Power cost matters more for the underlying AI compute buildout than for any single API call, since it's baked into what a provider charges rather than billed separately to developers. GPU hardware amortization is the bigger line item behind the scenes here. A single high-end AI accelerator can cost tens of thousands of dollars and typically depreciates over three to five years, so a provider's pricing has to cover that ongoing capital cost long before electricity ever becomes the dominant factor in the final bill.

A pen pointing to a financial graph showing sales and total costs.

How Do You Actually Budget for a Million Tokens?

Budgeting for a million tokens starts with estimating monthly usage volume, not a one-time cost, since most single tasks use only a tiny fraction of a million tokens. Multiply expected requests by average tokens per request, split input from output, and price each side separately against the model tier you actually plan to use.

Start with a worked example. A support-ticket summarizer processing 50,000 tickets a month, each averaging 600 input tokens and 150 output tokens, uses 30 million input tokens and 7.5 million output tokens monthly. On Claude Haiku 4.5, that's roughly $30 for input and $37.50 for output, about $67.50 a month. Route the same workload through a frontier reasoning model instead, and the output cost alone can climb past $200 for identical volume. That gap alone, roughly $67.50 versus $200-plus for the exact same 50,000 tickets, is usually a bigger lever than any single pricing negotiation. Most teams default to a stronger model out of caution during a pilot, then never revisit that choice once the workload scales to production volume.

Three levers cut real spend without touching model quality. Prompt caching, which several providers now offer, can cut repeated-context input costs by up to 90%. Batch processing, for workloads that don't need an instant response, typically runs at half the standard rate. And routing easy requests to a cheaper model while reserving frontier models for genuinely hard tasks, often called model routing, can cut blended costs 40% to 60% in practice.

The honest caveat: pricing changes fast. A rate published this month can shift within a quarter as providers respond to competition, so treat any specific number here as a September 2026 snapshot, not a fixed reference point a year out.

High-voltage electrical substation with steel structures and power lines at dawn.

The Bottom Line

A million tokens can cost anywhere from a few cents to over $50 depending on the model tier, and the gap keeps widening as reasoning models get more expensive while budget models get cheaper. The practical move for most teams isn't chasing the cheapest model across the board — it's matching each task to the cheapest model that still does the job, and watching usage volume as closely as the per-token rate. Model routing and prompt caching both genuinely help here, but the single highest-leverage move remains simple: audit which of your current production tasks are running on a model tier stronger than that specific task actually needs. Emergent Wire will keep tracking these numbers as providers adjust pricing through the rest of 2026, since a snapshot like this one has a shelf life measured in months, not years.

Emergent Wire covers AI models, capabilities, and the industry building them, for readers who want the numbers behind the headlines.

How much does a million tokens cost in 2026?
It depends heavily on the model tier. Budget models like GPT-5 nano or Claude Haiku 4.5 run well under $1 per million tokens combined, while frontier reasoning models like Claude Opus 5.5 or GPT-5.6-sol can cost $20 to $30 per million output tokens alone.
Why has AI inference gotten so much cheaper?
Model providers have combined smaller, more efficient architectures, better GPU utilization, and techniques like mixture-of-experts and quantization. Epoch AI has tracked equivalent-capability inference prices falling by roughly 50x per year on average since 2022.
Why is my company's AI bill going up if tokens are getting cheaper?
Per-token prices are falling, but usage is growing faster. Longer context windows, multi-step agent workflows, and reasoning models that generate extra 'thinking' tokens all multiply the number of tokens a single task consumes.
Does output token pricing matter more than input pricing?
Usually, yes. Output tokens typically cost 4 to 10 times more than input tokens across most providers, and reasoning models generate hidden output tokens during their internal reasoning steps, which adds to the output-side bill.
How does electricity cost factor into inference pricing?
Electricity is a real but usually smaller slice of the bill than GPU hardware amortization. Industrial electricity averaged roughly 8.5 to 8.8 cents per kilowatt-hour in the U.S. in 2026, according to the Energy Information Administration, though large data centers often negotiate different rates.