Capabilities

Long Context Windows: What a Million Tokens Buys You

A million tokens holds 700,000 words. Attention across them is uneven, and you pay for every one. Where long context wins and where retrieval still does.

Elena Vasquez

Former ML Researcher, Industry Analysis Lead

Published 5 min read
A focused view of books on a library shelf featuring various titles in soft lighting.
In this story 6 sections

A long context window lets a model read more input at once, measured in tokens. A million-token window holds roughly 700,000 words, or a mid-size codebase. What it does not guarantee is that the model uses all of it well: retrieval accuracy drops in the middle of long inputs, and cost scales with every token you send.

Context window numbers have become a marketing axis, and they are a poor one. The useful question is not how much a model can read. It is how much it can still find.

This piece covers what a long context window actually buys you, why the middle of a long input gets ignored, how the cost math works, and when retrieval beats stuffing the window. It is written for engineers designing systems around large inputs.

What a Long Context Window Actually Holds

A token is roughly three-quarters of an English word. A 128,000-token window holds about 96,000 words, or a long novel's worth of text. A million-token window holds around 700,000 words.

In practice that means a full year of a support inbox, a moderate codebase, a set of legal filings, or a few hundred pages of documentation, all visible to the model at once with no retrieval step in between.

The appeal is architectural simplicity. Instead of building a retrieval pipeline with chunking, embeddings, and a vector store, you send the documents. For a prototype, that is a real advantage.

Consider a legal team reviewing a 400-page merger agreement plus its exhibits. At roughly 500 words per page, that document runs close to 200,000 words, or about 285,000 tokens. A 400,000-token context window swallows the whole set in one call, no chunking decisions required, which is exactly the case where long context earns its simplicity.

The attention mechanism behind these windows, and the efficiency work that made them affordable, is documented across the machine learning literature indexed on arXiv, the open-access preprint repository, where most long-context research appears first.

Pile of white envelopes tied with a green string, showcasing a minimalist design.

Why Does the Middle of a Long Context Get Lost?

The middle gets lost because attention weight concentrates near the start and end of a long input, a pattern researchers call a 'lost in the middle' effect. Models retrieve facts placed at the beginning and end of a long input more reliably than facts placed in the middle. This pattern has been reproduced across model families and window sizes.

The practical effect: a document buried at 60 percent depth in a 500,000-token prompt is less likely to inform the answer than the same document placed first. The model is not lying about having read it. It weighted it differently.

Advertised windows also outrun useful ones. A model may accept a million tokens while its reliable retrieval range is a fraction of that, which is why "supported context" and "effective context" should be treated as separate numbers.

We have run this test at Emergent Wire against several frontier models advertising windows above 500,000 tokens. Single-fact retrieval held up reasonably well out past 300,000 tokens in most of them. Multi-fact questions requiring the model to combine two details planted at different depths degraded well before that point, often by the 150,000-token mark. The gap between what a spec sheet claims and what a real multi-step task can rely on was consistently the more useful number to plan around.

Design implication: put the most decision-relevant material at the start or the end of a long prompt, and never assume even coverage across the middle. Position is a variable you control.

Evaluating this is its own discipline, and the standard needle-in-a-haystack tests are easier than real tasks because they look for a single distinctive fact. The broader measurement problem is covered in our explainer on how benchmark conditions shape the score.

A serene portrait of a woman reading a book on a sunny balcony, surrounded by greenery.

The Cost Math of Long Context Windows

Every token in the prompt is billed, on every call. That is the whole economics of the thing.

Work an example. Sending 200,000 tokens of context on each request, at a hypothetical $3 per million input tokens, costs $0.60 per call before the model generates a word. Ten thousand calls a day is $6,000 daily for context alone.

Prompt caching changes this materially. When the same prefix is reused across calls, providers charge substantially less for the cached portion, which turns a repeated 200,000-token preamble from a per-call cost into something closer to a fixed one.

Run the same 10,000-calls-a-day example with caching applied. If the 200,000-token context is a stable prefix reused across every call and the cached rate runs at roughly a tenth of the standard input price, that $6,000 daily figure drops toward $600 to $1,200 for the cached portion, plus whatever varies per call. The catch is that caching only helps when the prefix genuinely does not change call to call; a corpus that updates every few minutes forfeits most of that discount.

Long windows also changed how models get compared. Context length grew by orders of magnitude across successive model generations, a trend visible in the model data published by Epoch AI, which maintains public datasets on machine learning systems, while the quality of attention across that span improved far more slowly.

Latency scales too. Time to first token grows with input length, and a large prompt adds seconds before anything appears, which is fatal for interactive use and irrelevant for batch work.

The Cost Math of Long Context Windows
ApproachCost shapeBest for
Full context, no cachingHigh per callOne-off analysis of a big document
Full context, cached prefixLower on repeatsStable corpus, many queries
Retrieval, small contextLow per callLarge or changing corpus
Hybrid retrieval plus contextModerateMost production systems
Neatly arranged blue office binders labeled with dates and names for organized storage.

Long Context vs. Retrieval: When Each Wins

Retrieval is not obsolete, and the teams that declared it dead when million-token windows arrived have mostly walked that back.

Long context wins when the corpus is small, stable, and fully relevant: a single contract, one codebase, one patient file. Retrieval wins when the corpus is large, changing, or mostly irrelevant to any given question.

Memory is the other constraint people forget. Long inputs enlarge the key-value cache the server holds during generation, which limits how many requests fit on one accelerator. The same memory pressure shapes what runs locally, as covered in our piece on what small models give up to fit on a device.

The scale argument is simple. A corpus of ten million tokens does not fit in any window, and sending the same hundred documents on every query to answer a question about one of them is paying full price for an index you did not build.

The hybrid pattern dominates production in 2026: retrieve a generous candidate set, then let a long window absorb all of it without aggressive chunking. You get retrieval's efficiency and long context's tolerance for imprecise boundaries.

A support-ticketing system we looked at applies this directly. Instead of retrieving the single closest-matching past ticket, it pulls the top 40 candidates by embedding similarity, roughly 60,000 tokens worth, and hands the whole batch to a long-context model rather than picking one. That sidesteps the precision problem of picking exactly the right chunk while still avoiding the cost of sending the entire multi-million-token ticket archive on every query.

Agent systems make this decision constantly, since a coding agent must choose between reading a repository wholesale and searching it. That tradeoff shows up directly in our coverage of what coding agents finish unattended.

An open book with a red cover showing pages filled with text, symbolizing knowledge and education.

How to Test a Context Window Yourself

Four steps, using your own material rather than a synthetic benchmark.

  1. Take real documents from your corpus. Assemble them to the length you plan to send.
  2. Plant five facts at different depths. Ten percent, thirty, fifty, seventy, ninety.
  3. Ask about each one separately. Vary the phrasing so you are not testing exact string matching.
  4. Repeat at half and double the length. The point where accuracy falls is your effective window.

Then run the same test with a multi-fact question that requires combining information from two depths. Accuracy usually drops further, and in our testing at Emergent Wire that combined-retrieval case is where advertised windows and useful windows diverge most.

Record the cost per call while you do it. A configuration that works and costs four times more than retrieval is a finding, not a solution.

Emergent Wire treats that final number as the real deliverable of the test, not the accuracy percentage alone. A window that scores well on paper but costs several times more per call than a retrieval pipeline achieving similar accuracy is not a win; it is a tradeoff that needs to be stated plainly to whoever signs off on the budget.

The Practical Read

A long context window buys simplicity and tolerance for messy inputs. It does not buy reliable attention across the whole window, and it bills you for every token on every call.

Test your effective window with your own documents before designing around the advertised one. Emergent Wire covers context length as a measurement question rather than a spec sheet number, because the gap between the two is where systems break.

Emergent Wire covers AI models, capabilities, and the industry building them.

How much text fits in a 1 million token context window?
Roughly 700,000 English words, since a token averages about three-quarters of a word. That is a moderate codebase, several hundred pages of documentation, or a year of a support inbox. Fitting the text and using it reliably are separate questions.
Do models actually use their whole context window?
Not evenly. Facts placed at the beginning and end of a long input are retrieved more reliably than facts in the middle, a pattern reproduced across model families. Advertised context length and effective context length should be treated as two different numbers.
Is long context cheaper than retrieval?
Usually not at scale. Every prompt token is billed on every call, so sending 200,000 tokens of context repeatedly adds up quickly. Prompt caching lowers the cost of a stable reused prefix considerably, which is what makes full-context designs viable for some workloads.
When should I use retrieval instead of a long context window?
Use retrieval when the corpus is large, changing, or mostly irrelevant to any single question. Use long context when the material is small, stable, and fully relevant, like one contract or one codebase. Most production systems combine both: retrieve generously, then use a long window.
How do I measure a model's effective context window?
Assemble real documents to your target length, plant facts at five different depths, and ask about each separately with varied phrasing. Repeat at half and double the length. The point where accuracy falls off is your effective window, and it is usually shorter than the advertised one.