The cycle repeats every few months with suspicious punctuality. A lab announces a larger context window, the announcement becomes a screenshot, and in the next meeting somebody says the line: «so we do not need the RAG anymore, we just paste the whole documents into the prompt». It sounds like simplification, which is what everyone wants to hear about an architecture that took months to build. And it is the kind of simplification you pay for three times: in accuracy, in billing, and at 2am when something answers wrong and nobody knows why.
The thesis in one line: «bigger context window or RAG» is not a fork in the road, because the two things do not solve the same problem. The context window is a capacity measure — how much fits in one call. Retrieval is a selection decision — what gets in, out of everything you have. Growing the first does not answer the second. And when you give up on deciding what gets in, what happens next is measured: accuracy falls well before the advertised limit, cost per call rises with the rate card in hand, and failure stops being debuggable because you no longer know which chunk the model used.
This is not about defending an architecture out of affection. There are cases — they are at the end, named — where the big window is genuinely the right answer and building retrieval would be over-engineering. It is about making the call for the real reasons rather than for a launch headline.
Bigger context window or RAG: not the same question asked twice
The confusion has a specific root: both things end up in the same place — text in front of the model at answer time — and that makes them look interchangeable. They are not. The context window is this call’s workspace: it fills up, gets used, and empties. Retrieval is the mechanism that picks, from a corpus that does not fit and never will, the chunks that deserve to occupy that space. One is the size of the table; the other is who decides which papers go on it.
The full taxonomy — conversation context, queryable knowledge and long-term memory, where each layer lives and what it costs — is already written in the guide to an AI agent’s memory, and there is no point re-arguing it here. What that guide does not cover, because it is not its question, is this one: what exactly happens when the market offers you a bigger table and you decide that saves you the person who picks the papers.
What breaks first is not the limit — it is accuracy
The counterintuitive part is that the problem shows up long before you fill the window. A model advertising a million tokens does not hold its quality until token 999,999 and then fall off a cliff: it degrades gradually, and fairly early. The NoLiMa benchmark measured this with a design that avoids the shortcut earlier «needle in a haystack» tests allowed — in NoLiMa the question and the relevant chunk share almost no wording, so the model cannot find it by literal match and has to infer the association, which is exactly what you ask of it in production.
The results: GPT-4o started at 99.3% with short context and dropped to 69.7% at 32K tokens and 56% at 128K. And it was not an outlier — at 32K, 11 of the 13 models evaluated fell to half or less of their own short-context baseline. Source: NoLiMa: Long-Context Evaluation Beyond Literal Matching, Modarressi et al., arXiv, February 2025.
On top of that sits a position effect documented earlier and independently: accuracy depends on where the information sits inside the context. The paper that coined the term describes a U-shaped curve — the model recovers what is at the start and the end reasonably well, and fails on what is left in the middle. Source: Lost in the Middle: How Language Models Use Long Contexts, Liu et al., Transactions of the ACL, 2024.
Together, the two effects describe the real failure mode of the «just dump it all in» strategy. It is not that the model refuses to answer: it is that it answers confidently using what it found, which is not necessarily what mattered. The answer arrives well written, well structured and badly grounded. That is the worst kind of error an enterprise system can have, because it is indistinguishable from a correct one without checking it by hand.
| What the announcement says | What the benchmark measures | What it means for your system |
|---|---|---|
| «1M token window» | Degradation starts well below that number | The advertised limit is input capacity, not a quality guarantee |
| «Your whole documentation fits» | At 32K, 11 of 13 models fall to ≤50% of baseline (NoLiMa, 2025) | Fitting is not the same as being used |
| «The model finds what it needs» | What sits mid-context is recovered worse (Liu et al., 2024) | The order you stack documents in becomes a hidden variable |
| «You skip building retrieval» | Benchmarks measure neither the bill nor traceability | Saved engineering turns into recurring cost and opacity |
The vendor selling you the million tokens charges double for them
No argument needed here: just read the rate card of the company selling the big window. In the public Gemini API price list, two models have their price split by prompt length, with the cut at exactly 200,000 tokens. On Gemini 3.1 Pro Preview, standard tier, input goes from $2.00 to $4.00 per million tokens once you cross that threshold, and output from $12.00 to $18.00. On Gemini 2.5 Pro, input goes from $1.25 to $2.50 and output from $10.00 to $15.00. Context caching doubles too. Source: Gemini Developer API pricing, Google, accessed 4 October 2026.
What matters is not the amount, which will change. It is the shape of the rate card: the vendor offering you the enormous window has decided to charge you double for actually using it. That is a statement about the real cost of serving those prompts, written by the party that pays it. When somebody frames «bigger context window or RAG» as if the first option were the free one, they are ignoring that the manufacturer itself has priced it as a different, more expensive product.
And the compounding is what wrecks budgets. Retrieval concentrates spend on building and maintaining an index: a cost that happens once and amortises across every query. Dumping everything into the prompt moves that spend to the variable side, where it multiplies by every call, every day, forever. A pilot with two hundred queries a month will not notice. The same system opened to three hundred people will. The arithmetic is boring, which is why nobody runs it before the meeting: multiply your input tokens by your expected volume and compare it with what an index costs. The guide on which model to use in an agent covers the rest of that decision, because window size is one spec among several and almost never the deciding one.
The failure you cannot debug
This is the argument that shows up in no benchmark and costs the most in operation. With retrieval, when an answer comes out wrong you have a trail: you know which chunks were retrieved, with what query, at what score. Diagnosis reduces to two answerable questions — was the right chunk in the index? did retrieval surface it? — and each points at a different fix: re-ingest the source, or tune retrieval.
Without retrieval there is no trail, because there was no selection to log. You handed over two hundred thousand tokens and the model used what it used. You cannot know what it leaned on, you cannot reproduce the path, and you cannot fix it with a scoped intervention: your only lever is reordering documents and trying again, which is debugging by superstition. In an internal system with high tolerance, that is an annoyance. In a system that answers customers or feeds a decision, it is the difference between an incident that closes and one that stays open.
When the big window is the right answer
Building retrieval over a corpus that does not need it is the other way to get this wrong, and it is more common than it sounds. Three situations where the big window wins cleanly:
- The corpus is small and stable. If everything the system needs to consult fits comfortably below the threshold where the model degrades, and it does not change weekly, an index is infrastructure you maintain for no gain.
- The evidence is spread across the whole document. When the task is synthesising, comparing sections or spotting contradictions across a long text, retrieval works against you: chunking is precisely what destroys the relationship between the parts. Here long context is not a luxury, it is the requirement.
- It is a one-off analysis, not a system. A contract, a report, a data dump analysed once. Nobody should build an ingestion pipeline for a question asked a single time.
The honest reading of the recent literature is that the binary framing is outdated from both ends: long context performs better when evidence is distributed, and retrieval performs better when evidence is sparse and has to be found. The architecture that is winning does not choose: it retrieves to narrow the universe down to the plausible, then uses the long window to reason over that reduced set. Retrieval stops being a surgical precision filter and becomes a noise reducer, which relaxes the requirements on your index considerably.
The two-minute test before you rip anything out
If somebody on your team proposes removing retrieval because a model shipped with more context, these four questions settle the conversation without a follow-up meeting:
- How many tokens is the full corpus, today and in twelve months? If the twelve-month answer crosses the threshold where the model degrades — and that sits well below the advertised limit — the window is not a solution, it is an extension.
- What would the real query volume cost at long-prompt rates? With the price split at 200K tokens, the maths fits on one sheet. Run it at the volume you are aiming for, not the pilot’s.
- Is the typical evidence in one place or spread out? Specific and localised: retrieval. Distributed across the document: long context. Both depending on the question: hybrid, and that is the most frequent answer.
- What do you say when a customer asks where that sentence came from? If the answer has to be verifiable, you need the record of what went in. Only selection gives you that.
None of the four is answered by window size, which is exactly the point.
What this changes in operations
There is a less technical reason the «drop the RAG» pitch lands so well, and it is worth naming: it is not that the big window is better, it is that maintaining an index is continuous work and nobody fancies doing it. Knowledge ages, sources change, documents get replaced, and if nobody re-ingests or invalidates what expired, the system starts answering with last year’s version. That is the real crack the infinite-context argument exploits: it promises to get rid of an operating function, not a piece of software.
The problem is that the function does not disappear, it just goes invisible. If you paste whole documents into the prompt, you still need those documents to be the current ones — except now you have neither an index nor a pipeline where you could check. Which is why knowledge-base freshness is a continuous function you operate, not a project you close: per-source refresh, invalidating what expired and an owner for every editorial call exist either way, with RAG or without it.
And if the argument you actually have open is not this one but the other — whether to retrain the model on your data — it is settled elsewhere and on the same logic: retrieval beats fine-tuning almost every time, on cost and on the ability to update without retraining. The three conversations — window, retrieval and training — get mixed into the same meetings, and it pays to keep them apart, because only one of them changes when a new model ships.
The operating conclusion is short. More context is good news: it lets you pass more relevant evidence per call and relaxes the precision you demand of your retrieval. What it does not do is decide what is relevant. That work still gets done by somebody — an index, a query, a policy — or it gets done by nobody, and then what you have is not a simpler architecture: it is the same complexity, with no record, at double the rate.