Prompt engineering asks "how do I phrase the instructions?". Context engineering asks a broader question: of everything the model could see at this step, what should it see? For a single chat completion the difference is small. For an agent that runs for fifty steps, reads files, calls tools and accumulates results, it is the difference between a system that stays sharp and one that slowly loses the thread. This guide explains why context is a scarce resource, and the techniques that keep it focused.
Why "just use the big context window" fails
Modern models accept hundreds of thousands or even a million tokens. It is tempting to put everything in: the whole codebase, all documentation, the full conversation history. In practice, quality degrades well before the limit:
- Attention is not uniform. Models retrieve information from the beginning and end of long contexts better than from the middle (Liu et al., Lost in the Middle).
- Effective context is shorter than advertised. Benchmarks that go beyond simple "needle in a haystack" retrieval show performance dropping as contexts grow, especially for multi-hop reasoning (Hsieh et al., RULER).
- Distractors hurt. Irrelevant but plausible content leads models astray more than missing content does.
- Cost and latency scale with tokens. Every step of an agent re-reads the whole context; see LLM cost optimization.
Anthropic's engineering team describes context as a finite resource with diminishing marginal returns — an "attention budget" (Effective context engineering for AI agents). The goal is the smallest set of high-signal tokens that maximises the chance of the desired outcome.
What competes for the context window
| Component | Typical size | Notes |
|---|---|---|
| System prompt | 1–5k tokens | Stable; cacheable |
| Tool definitions | 0.5–20k tokens | Grows with every tool; see tool design |
| Retrieved documents | 2–50k tokens | Biggest lever in RAG |
| Conversation history | Grows unbounded | Needs trimming or summarising |
| Tool results | Grows every step | Often the largest consumer in agents |
| Memory / notes | 0.5–5k tokens | Curated facts from earlier work |
| Current task | Small | Should never be pushed out |
Treat this like a memory budget in embedded programming: every component gets an allowance, and something has to justify every token.
Technique 1: Right-size the system prompt
System prompts fail in two directions. Too specific — long lists of if-then rules for edge cases — and they become brittle and contradictory. Too vague — "be helpful" — and the model lacks the signals it needs. Aim for the middle: clear principles, the few hard rules that matter, and a handful of canonical examples rather than an exhaustive list. Organise with headings or XML tags so sections are easy to locate. The details are in prompt engineering for developers.
Technique 2: Just-in-time retrieval instead of preloading
Instead of loading all potentially relevant data upfront, give the agent lightweight references and tools to fetch details when needed: file paths, document IDs, search tools, database queries. This is how coding agents work in large repositories: they do not read every file; they grep, list directories, open the relevant files and keep only what matters.
Trade-offs:
- Preloading is faster at runtime and works when the relevant set is small and predictable (a customer's profile for a support chat).
- Just-in-time scales to large corpora and keeps context clean, at the cost of extra tool calls.
- Hybrid — preload a small, high-value core (instructions, user profile, a project instruction file like
CLAUDE.md), retrieve the rest on demand — is the default we recommend.
Retrieval quality then becomes critical; see RAG for business and embeddings and chunking.
Technique 3: Trim and summarise tool results
Tool results are the fastest-growing part of an agent's context. Strategies:
- Shape outputs at the source. Tools should return concise, high-signal results with pagination; this is cheaper than cleaning up afterwards.
- Clear stale results. Once a large tool output has been used — a file read, a search result — it rarely needs to stay verbatim. Some platforms support automatic clearing of old tool results; otherwise replace them with a one-line note ("read config.yaml: uses Postgres 16, pool size 20").
- Store large artefacts outside the context. Write big outputs to files and give the agent the path.
Technique 4: Compaction
When a conversation approaches the context limit, summarise it and continue in a fresh window with the summary. Done well, compaction keeps decisions, open problems, important details (file names, IDs, error messages) and discards redundant tool output and dead ends.
Summarise this session for continuation. Keep:
- the task and acceptance criteria
- decisions made and why
- files changed and their current state
- unresolved errors with exact messages
- next planned steps
Drop: raw tool outputs already acted on, failed approaches (one line each).
Tune the summary prompt on real long sessions: first maximise recall (nothing important lost), then trim. Coding agents implement this as automatic or manual compaction.
Technique 5: Structured note-taking
For long tasks, let the agent maintain notes outside the context window — a NOTES.md file, a to-do list, a memory tool — and re-read them as needed. Notes survive compaction and context resets, and they give the agent a persistent sense of progress across hours of work. This works especially well for multi-step migrations, research tasks and anything that spans multiple sessions.
# NOTES.md — migration of billing to java.time
- [x] Package billing.invoice (12 files) — tests green
- [x] Package billing.tax — note: TaxPeriod uses Europe/Kyiv explicitly
- [ ] Package billing.reports — 3 usages of Joda Interval, need adapter
Open question: rounding in ReportTotals differs from old behaviour? (#481)
Technique 6: Subagents with clean contexts
A subagent is a separate model call or agent with its own fresh context, given a focused task — "find all places where we parse dates in the reports package and summarise the formats used" — and returning only a condensed result. The main agent's context receives a few hundred tokens instead of the tens of thousands the subagent consumed.
This separation of concerns is one of the main reasons multi-agent architectures outperform single agents on broad research tasks, at a significant token cost; see multi-agent systems. Coding tools expose the same idea as subagents.
Technique 7: Long-term memory, carefully
Memory across sessions — user preferences, past decisions, facts about a project — makes agents feel smarter, but it is also context that can be wrong, stale or injected. Rules we follow:
- Store facts with provenance and dates, not raw transcripts.
- Retrieve memory selectively by relevance, like any other document.
- Let users see and delete what is remembered about them — this is also a privacy requirement.
- Treat memory written from untrusted content as untrusted.
Ordering and caching
How you arrange context matters for both quality and cost:
- Stable content first: system prompt, tool definitions, long reference documents. This prefix can be cached, which cuts cost and latency substantially for repeated calls.
- Semi-stable next: user profile, project notes.
- Dynamic last: conversation history, latest tool results, the current question.
Placing long documents before the question also tends to improve answer quality.
Measuring context quality
Context engineering is testable:
- Ablation: remove a context component and re-run evals; if scores do not drop, you were paying for nothing.
- Token budget per step: track input tokens per agent step over a session; sudden growth points to bloated tool results.
- Long-session evals: include tasks that require 30+ steps so that degradation shows up in tests, not in production.
- Trace review: read what the model actually saw at the step where it went wrong; see LLM observability.
FAQ
Is context engineering just a new name for RAG? RAG is one technique within it. Context engineering also covers instructions, tools, history, memory and how all of them evolve during a multi-step task.
Do bigger context windows make this obsolete? No. Larger windows raise the ceiling, but focus and cost still depend on what you put in. Every published long-context benchmark shows degradation with more distracting content.
How do we decide what to summarise? Keep anything needed to make future decisions — goals, constraints, decisions, open errors, identifiers — and drop what has already been acted on.
Where should we start? Measure tokens per step in your agent, find the largest contributor (usually tool results or history), and fix that first.
Sources
- Anthropic (2025). Effective context engineering for AI agents.
- Liu et al. (2023). Lost in the Middle: How Language Models Use Long Contexts.
- Hsieh et al. (2024). RULER: What's the Real Context Size of Your Long-Context Language Models?
- Anthropic. Prompt caching.
- Anthropic. Claude Code subagents.
- Packer et al. (2023). MemGPT: Towards LLMs as Operating Systems.