The Real Choice Is What The Task May See
A context window tells you how much a model can accept. It does not tell you what a business task should receive.
At one pole, you provide the whole relevant corpus and let the model identify what matters. At the other, you maintain small source files, explicit routing, and one owner for every important fact.
Both methods allocate work differently. Broad context asks the model to filter during the run. Hard curation asks people and systems to filter before or during retrieval.
Steps
Guide
Separate Context Capacity From Context Choice
A model may accept a long input without using every part equally well. Maximum context length is capacity; reliable use is a task result.
The Lost in the Middle research tested multi-document question answering and key-value retrieval. In those experiments, performance often fell when relevant information appeared in the middle of a long input rather than near the beginning or end.
That does not prove every current model fails on every long task. It gives you a strong reason to test the position and volume of important information instead of trusting a context-window number alone.
Use Karpathy's Actual Context Rule
Andrej Karpathy helped popularise the term context engineering. In his own public post, he described the goal as providing just the right information for the next step. He included memory, tool outputs, search results, and instructions in that system.
This is a system-design argument, not a literal instruction to load every available file into every run.
The useful question is therefore not whether the model is clever enough to filter everything. Ask what information this step needs, how it becomes available, and what should remain outside.
Know When Broad Context Earns Its Cost
Broad context reduces preparation. The model can inspect a bounded corpus and notice connections a curator did not predict. This can help exploratory work, especially when the material changes slowly and omissions cost more than extra review.
The cost is not only window space. Many APIs meter input tokens, while prompt caching may reduce the price of repeated prefixes; Anthropic's API pricing shows both as separate usage categories. Retrieval and tool calls also cost time and money, so compare the full run rather than one prompt length.
Broad context becomes risky when the corpus contains duplicate facts, several policy versions, or material with unclear authority. Visibility is not ownership. A fluent answer can still come from the stalest visible copy.
Know When Hard Curation Pays Back
Hard curation gives every important fact one owner file. A thin routing file tells the agent where to look, which material is reference-only, and when a conflict requires human review.
This is useful when work repeats, several people or agents use the same material, or an incorrect answer has serious consequences. The team can preserve the files and versions used for a run.
The cost is maintenance. A clean folder tree does not stay correct by itself. A stale router can reproduce the same wrong route every time. Give routing entries owners and update triggers.
Compare The Five Decisions That Matter
Do not choose a method from taste alone. Compare it against the operating conditions of the task:
- Context size and per-run cost: count the whole workflow, including retrieval, tools, cache behaviour, and review time.
- Change frequency: rapidly changing sources need clear ownership and routes; stable bounded corpora tolerate broader loading.
- Reproducibility: preserve the exact corpus snapshot or the exact routed files used.
- Shared use: more readers and agents increase the value of explicit source authority.
- Consequence of error: higher stakes justify stronger traceability, stop conditions, and human review.
A method that wins for exploratory research may lose for a recurring approval. Choose by job, not by person or vendor.
Build The Practical Middle
Most durable setups combine a small, always-loaded router with material fetched on demand. The router carries rules and locations, not the whole knowledge base.
It should state:
* which source owns each type of fact
* which files are instructions and which are reference-only
* what the agent may never change
* when wider search is allowed
* when conflicting material must go to a personAnthropic describes a similar hybrid in its context-engineering guidance: instruction files can enter up front, while file-navigation tools retrieve other material just in time.
Retrieval implements the middle; it is not a third philosophy. Sub-agents do the same at another scale by exploring focused areas in separate contexts and returning distilled findings. Anthropic documents that pattern in its multi-agent research system.
Run The Same Task Both Ways
Use one task, one source snapshot, one model, one configuration, one prompt, and one review standard. Change only the context method.
For Version A, load the broad relevant corpus. For Version B, load the thin router and only the material it selects. Record input usage, latency, output, sources used, unsupported claims, and reviewer corrections.
Review the outputs without labelling which method produced them. Repeat the test when model variation could affect the conclusion. One attractive answer is not a durable policy.
This moves the choice from argument to evidence. It also fits the wider move from isolated prompting to persistent AI project context and workflow orchestration.
Write A Policy By Task Type
The output is not one universal winner. It is a short rule that another person or agent can apply without repeating the debate.
Use broad context when exploration matters, the corpus is bounded, and omission is more costly than reviewing extra material. Use harder curation when work repeats, sources change often, several readers depend on them, or errors carry serious consequences.
Use the hybrid when you need a stable operating frame and room to investigate. Maintain the source files as a reusable knowledge system, not as a growing pile of copied context.
The durable advantage is knowing which context produced the answer, why it was selected, and how the rule will stay current.
Be Aware
The broad-context test contains several versions of the same fact.
Mark the authoritative owner and archive status before testing. The comparison should test context method, not unresolved source governance.
The curated version looks clean but misses an unanticipated source.
Allow a bounded wider search, log the miss, and add a permanent route only when real use shows it is needed.
The two versions use different prompts or source dates.
Freeze every variable except the context method, then rerun the pair.
Sub-agents repeat the same work or return incompatible notes.
Give each sub-agent a distinct boundary, source type, and output format before delegation.
The team declares a universal winner after one run.
Repeat consequential tests and write the decision for that task type, not every possible AI job.
Context Method Comparison
Copy / paste
Compare broad context with routed context for one real task.
#### Freeze The Test
* Task:
* Expected output:
* Source snapshot and date:
* Model and configuration:
* Shared task prompt:
* Acceptance criteria:
* Consequence if the answer is wrong:
#### Version A: Broad Context
* Material loaded:
* Input usage:
* Retrieval or tool calls:
* Latency:
* Available cost data:
* Sources cited or referenced:
* Correct statements:
* Omissions:
* Unsupported or stale statements:
* Human corrections required:
#### Version B: Routed Context
* Always-loaded router:
* Material fetched:
* Input usage:
* Retrieval or tool calls:
* Latency:
* Available cost data:
* Sources cited or referenced:
* Correct statements:
* Omissions:
* Unsupported or stale statements:
* Human corrections required:
#### Decision
* Winning method for this task type:
* What is always loaded:
* What is fetched on demand:
* When wider search is allowed:
* What is excluded:
* Human review or stop condition:
* Owner of the routing rule:
* Review trigger:
Keep the source snapshot, model, configuration, prompt, and acceptance criteria identical. If any of those change, stop and restart the comparison. Do not invent missing cost or quality measurements.



