AI

Context Engineering for LLM Agents

Agents fail when their context fills with noise. Get the write, select, compress, isolate playbook, real token data, and a free planner to cut cost and errors.
Context Engineering for LLM Agents

Introduction

Context engineering for LLM agents is the craft of deciding exactly what a model sees at every step of a long task. A chat prompt is read once, but an agent loop rereads a growing pile of instructions, tool outputs, and history on every turn. The team behind Manus reports an average input-to-output token ratio of 100 to 1 in production, so agents spend nearly all their effort reading. Research from Chroma tested 18 language models and found that performance grows increasingly unreliable as input length grows. A bigger window is therefore not a free upgrade, because every extra token competes for the same limited attention. This guide walks through the full toolkit, from writing notes outside the window to compressing, selecting, and isolating context across sub-agents. It also covers caching economics, security, ethics, measurement, and the road ahead, with an interactive planner you can use to size your own agent.

Quick Answers on Context Engineering for LLM Agents

What is context engineering for LLM agents?

Context engineering for LLM agents is the practice of curating the smallest set of high-signal tokens, including instructions, tools, memory, and retrieved data, that gives an agent the best chance of completing each step correctly.

How is context engineering different from prompt engineering?

Prompt engineering writes the instructions, while context engineering manages everything the model sees across many turns, including tool results, memory, and history. It treats the context window as a budget to curate, not a page to fill.

What are the main context engineering strategies for agents?

Four strategies dominate: write context to external memory, select relevant context through retrieval, compress context with summaries or trimming, and isolate context across sub-agents or sandboxes so each one stays small and focused.

Key Takeaways

  • Context is a finite budget: every token you add depletes the model’s attention, so curate the smallest high-signal set rather than filling the window.
  • Four moves cover most needs: write context out to memory, select it back in on demand, compress it as it ages, and isolate it across sub-agents.
  • Tool definitions, tool outputs, and cache-hostile prompt layouts quietly drive most of the cost and failure in production agents.
  • Treat context quality as an engineering metric with evals, traces, and owners, because longer windows alone do not prevent degradation.

What Is Context Engineering for LLM Agents in Plain Terms

Context engineering for LLM agents is the practice of curating and managing the tokens an agent sees at each step, including instructions, tools, memory, and retrieved data, so the model stays accurate, efficient, and on task across long, multi-turn work.

An Interactive From AIplusInfo

Context Budget Planner for an LLM Agent

Change the window, the tool catalog, and the management strategy to see how full the agent’s context gets and what repeated reading costs.

200,000 tokens

Smaller windowLarger window

No management

SimplestMost structured

25 tools

0 tools80 tools

1,500 tokens

200 tokens8,000 tokens

50 calls

5 calls150 calls

Peak live context

0

0 percent of the window

Window health

Comfortable

Overflow check pending

Input cost, no cache

$0.00

0 input tokens billed

Input cost, cached prefix

$0.00

0 percent saved

Black: system promptGold: tool definitionsBlue: working historyGrey: free space

Adjust a control to see a recommendation.

Illustrative model. Assumes 2,500 tokens of system prompt, 350 tokens per tool definition, 250 tokens of reasoning per call, and prices of $3.00 uncached and $0.30 cached per million input tokens, the Claude Sonnet figures reported in the Manus write-up. Quality often slips well before a window fills; Databricks results cited by Drew Breunig showed correctness falling near 32,000 tokens for a large open model.

From Prompt Craft to Context Engineering

Prompt engineering earned its reputation in the era of single-turn chat, when a clever phrasing could noticeably change an answer. That craft still matters, but an agent is a loop rather than a question, and the loop keeps generating new material that lands back in front of the model. Andrej Karpathy popularized the newer framing by describing the job as filling the context window with just the right information for the next step. The shift in vocabulary reflects a real shift in where the leverage lives, because the instruction text is now a small slice of what the model reads. A system prompt might run a few thousand tokens, while tool results and history can grow to a hundred thousand during a single session.

Anthropic describes the same idea as asking what configuration of context is most likely to generate the behavior you want. In its guidance on effective context engineering, prompt engineering becomes one component inside a wider system that also manages tools, memory, and retrieved data. Everyday usage agrees, and our look at why shorter prompts improve accuracy covers a Meta AI study where concise inputs lifted results by up to 34 percent. Fewer, better tokens consistently beat more tokens, and that principle scales from a single prompt to an entire agent run. Teams that internalize it stop asking how to word instructions and start asking what the model should be allowed to see.

The Anatomy of an Agent's Context Window

Stepping back from the buzzword, an agent's context window is a bundle of distinct ingredients that compete for space. The system prompt sets role, rules, and tone, and it usually stays constant for the life of the session. Tool definitions come next, and each function or server you connect adds names, descriptions, and parameter schemas that are re-read on every turn. The message history then accumulates the user requests, the model's own reasoning, and every tool call with its result. Retrieved documents, memory snippets, and scratchpad notes are added on top when the agent decides it needs them.

Each ingredient has a different lifetime and a different cost profile, which is why it helps to sort them by how often they change. Stable content such as the system prompt and tool schemas belongs at the front of the window, where it can be cached and reused cheaply. Volatile content such as search results and file contents belongs near the end, where it can be replaced without invalidating everything before it. Anthropic's tooling guidance notes that Claude Code caps tool responses at 25,000 tokens by default. That blunt limit is a reminder that one careless tool call can swallow a large share of the budget. Once you can name each ingredient and its owner, you can start deciding which ones deserve their tokens.

Two numbers frame the practical limits of any agent design. The first is the hard ceiling, such as the 200,000 tokens at which Anthropic's research agent had its context truncated and had to save its plan to memory. The second is the soft ceiling, the point where quality starts to slip long before the window is full. Most agent failures live between those two lines, in the zone where the prompt still fits but the model no longer uses it well. Good context engineering is largely the work of keeping the live window inside that comfortable middle.

Why Bigger Windows Do Not Fix the Problem

Beyond the headline token counts, the research on long inputs tells a consistent and slightly uncomfortable story. In the paper Lost in the Middle, Liu and colleagues showed that accuracy peaks when relevant facts sit at the start or end of the input. Accuracy drops substantially when those same facts sit in the middle. Chroma's later study of 18 models found that even simple tasks become less reliable as inputs lengthen, and that a single distractor can reduce accuracy. Anthropic frames the cause as an attention budget, because every token attends to every other token and the resulting n squared pairwise relationships grow strained as context lengthens. The model does not read a long prompt the way a person skims a page; it spreads a finite resource across everything you give it.

The practical consequence is that advertised window sizes overstate usable context. Chroma's LongMemEval experiment compared a focused input of roughly 300 tokens with a full input of about 113,000 tokens. The focused version produced clearly higher accuracy across model families. Our deep dive on context rot in large language models covers the measurement side in detail, so here the takeaway is simply about design. If a task can be answered from 300 well-chosen tokens, sending 113,000 is not generosity, because it handicaps the model you are trying to help. Every unnecessary block also raises the bill, since agents pay for those tokens again on every turn of the loop.

Longer windows still have real value, because some tasks genuinely need large inputs such as whole codebases or long contracts. The point is not that big windows are useless but that they should be treated as capacity you spend deliberately. A sensible rule is to ask, for every block of text, what decision it helps the model make on the next step. Blocks that cannot answer that question are candidates for removal, summarization, or storage outside the window. That habit, applied consistently, is most of what context engineering means in day-to-day practice.

How Context Goes Wrong: Poisoning, Distraction, Confusion, and Clash

Looking at failures in the wild, Drew Breunig's taxonomy of context failures gives teams a shared vocabulary. Context poisoning happens when a hallucination or error enters the context and is then referenced again and again. His example is Gemini 2.5 playing Pokemon, where a corrupted goals section could leave the model fixated on impossible or irrelevant objectives. The fix is to validate what gets written to long-lived context and to keep a clean path for quarantining or rolling back bad entries. Poisoning is dangerous precisely because the model trusts its own history more than it should.

Context distraction is the second mode, in which accumulated history pulls the model toward repeating past actions instead of reasoning freshly. In the same Gemini experiment, once the context grew well past 100,000 tokens the agent showed a tendency to favor repeating actions from its history rather than synthesizing new plans. Databricks research cited in the same piece found that correctness for Llama 3.1 405B began to fall around 32,000 tokens, and earlier for smaller models. The lesson is that the effective window of many models is far smaller than the number on the box, so compaction should start early rather than at the limit. Distraction is a signal to summarize, not a reason to buy a bigger model.

Context confusion arises when superfluous content, especially extra tools, leads the model to produce worse responses. The Berkeley Function-Calling Leaderboard showed that every model performs worse when given more than one tool. A GeoEngine test found that a quantized Llama 3.1 8B failed with all 46 tools but succeeded with a curated set of 19. Context clash, the fourth mode, appears when new information contradicts earlier turns, as in research where information spread across multiple turns produced an average 39 percent drop in performance. Both problems argue for fewer, sharper tools and for consolidating facts into one authoritative statement rather than leaving contradictions scattered through the history. For a wider view of how agent teams manage these hazards, see our guide to deterministic guardrails for AI agents.

Each failure mode maps to a defensive habit, and the habits compound when used together. Validate writes to guard against poisoning, summarize early to resist distraction, prune tools to prevent confusion, and reconcile facts to avoid clash. None of these requires a new model or a new vendor, only discipline in how the context is assembled. Teams that keep a short list of these checks in their code review process catch most problems before users do. The remaining sections turn each habit into a concrete technique.

Writing Context to Memory Outside the Window

Turning to the first of the four strategies, writing means saving information outside the context window so the agent can retrieve it later. The simplest form is a scratchpad, where the agent records a plan, a checklist, or intermediate findings through a tool call that writes to a file or a state object. Anthropic's research system uses exactly this pattern, saving its plan to memory because the context could be truncated after 200,000 tokens. Writing converts a perishable resource, the window, into a durable one, the file system or database. An agent that can write down what it learned no longer has to keep every detail alive in the live prompt.

Memories extend the idea across sessions rather than within a single run, so the agent can improve with use. Without them, every conversation starts from zero and the user must repeat preferences, constraints, and background facts. Approaches such as Reflexion and Generative Agents let a system synthesize memories from past interactions. Products like ChatGPT, Cursor, and Windsurf now keep user-level memory between conversations, as LangChain's overview notes. Anthropic's long-running harness work shows a practical variant of the same idea. An initializer agent sets up a progress file and a feature list of over 200 items, so every new session can pick up where the last one stopped. Our explainer on agent memory architecture covers working, episodic, semantic, and procedural layers, so this section stays focused on how writing fits the wider budget.

Selecting the Right Context at the Right Moment

Building on that foundation, selection is the mirror image of writing, because it decides which stored material returns to the window. The oldest form is retrieval-augmented generation, which embeds documents and fetches the closest matches for each query. Choosing between retrieval and training is itself a design decision, and our comparison of retrieval-augmented generation versus fine-tuning lays out when each one pays off. For corpora with rich relationships, a graph can outperform plain vector search, which is the subject of our piece on GraphRAG compared with traditional RAG. The core discipline is returning less, not more, because every retrieved passage spends attention the model could have used elsewhere.

Agents add a twist that classic retrieval pipelines did not face, which is that the agent can decide what to fetch. Anthropic calls the pattern just-in-time retrieval, where the agent keeps lightweight identifiers such as file paths, URLs, or stored queries and loads data only when needed. This mirrors how people work, since nobody memorizes an entire library but everyone knows how to find the right shelf. Metadata like folder names, naming conventions, and timestamps give the agent cheap signals about what is worth opening. The tradeoff is speed, because runtime exploration is slower than a precomputed index, and the agent can wander into dead ends without good tool design.

Tool selection deserves the same care that teams already give to document selection. LangChain reports that applying retrieval to tool descriptions can improve tool selection accuracy by roughly threefold when an agent has many tools. Instead of loading fifty schemas into every prompt, the agent searches a catalog and loads only the two or three it needs. Code agents show the most advanced version of this approach, combining grep, semantic search, knowledge graphs, and re-ranking to find the right file. Structured relationships make that search more precise, because the agent can follow links between entities instead of guessing from raw similarity.

Ordering and placement matter as much as the choice itself. Because of the lost-in-the-middle effect, the most important instructions and facts should sit near the beginning or the end of the window rather than buried in the center. Manus uses a related trick by continually rewriting a todo file so the current objectives stay in the most recent tokens, a technique it calls recitation. Coding agents benefit from live reference material, since current specifications reduce integration bugs that stale documentation would otherwise cause. Selection is therefore two decisions at once, what to include and where to put it.

Compressing Context With Summaries and Compaction

Moving on to compression, the goal is to retain only the tokens required to perform the task at hand. Compaction is the best known technique, and it works by summarizing a conversation that is nearing its limit and restarting the loop with that summary. Anthropic describes how Claude Code preserves architectural decisions, unresolved bugs, and implementation details while discarding redundant tool outputs and messages. The recommended tuning approach is to maximize recall first so nothing critical is lost, then iterate on precision by cutting what turned out to be superfluous. According to LangChain's overview, Claude Code triggers its auto-compact once roughly 95 percent of the window is used, which is later than many teams should wait.

Trimming is the blunter cousin of summarization, and it is often the cheaper choice. Hard-coded heuristics can drop the oldest messages, and trained pruners such as Provence can filter out irrelevant passages from retrieved text. Anthropic's context editing feature takes a more surgical route by automatically clearing stale tool calls and results as the window fills, while leaving the conversation flow intact. Old tool output is the best candidate for removal because the agent already acted on it, and the result rarely needs to be reread verbatim. Clearing it while keeping the agent's own conclusions preserves the reasoning trail at a fraction of the token cost.

Compression always trades fidelity for space, so it needs guardrails to avoid erasing the wrong things. The Manus team compresses observations in a restorable way, keeping a URL or file path in context even after the full page or file contents are dropped. If the agent needs the details again, it can fetch them, which turns a lossy summary into a reversible one. Failed attempts deserve special care, because the same team found that leaving errors and stack traces visible helps the model avoid repeating them. A summarizer that tidies away every failure produces a cleaner prompt and a worse agent, so instruct it to keep decisions, open problems, and dead ends.

Isolating Context With Sub-Agents and Sandboxes

Beyond single-agent loops, isolation means splitting context across components so that each one sees only what it needs. A sub-agent receives a focused task and a clean window, explores freely, and returns a condensed summary of its findings. Anthropic notes that such sub-agents may use tens of thousands of tokens while exploring yet hand back summaries of only 1,000 to 2,000 tokens. The lead agent therefore keeps a small, high-signal context while the heavy reading happens elsewhere, a pattern explored further in our guide to hierarchical coordination in multi-agent tasks. Sandboxes extend the same logic to bulky objects, because a code environment can hold images, audio, or large tables that the model never needs to read directly.

Isolation carries a cost that many teams discover only after the first demo. Cognition's essay on why not to build multi-agent systems argues that agents should share full traces, not just messages. Actions carry implicit decisions, and those decisions conflict when parallel workers cannot see each other's work. Its Flappy Bird example has one subagent drawing a Super Mario style background while another builds an incompatible bird. Isolation also multiplies spend, since Anthropic measured multi-agent research at roughly 15 times the tokens of a chat interaction. The sensible rule is to isolate read-heavy, parallel, loosely coupled work, and to keep tightly coupled decisions inside one shared context.

Designing Tools That Respect the Token Budget

Shifting attention to tools, every function you expose is context, and it is paid for on every turn. Names, descriptions, and parameter schemas sit in the prompt whether or not the agent uses them, so a large catalog quietly taxes every request. Anthropic's guidance puts the test plainly: if a human engineer cannot say which tool fits a situation, the agent cannot be expected to do better. Overlapping tools create exactly that ambiguity, which is the context confusion described earlier. Our primer on function calling in LLMs shows how schemas are defined and why each field matters for selection.

Tool outputs are the second budget item, and they often dwarf the definitions. The same Anthropic tooling guidance describes a response format option that lets the agent choose concise or detailed results. In one Slack example the concise form used about a third of the tokens, 72 versus 206. Pagination, filtering, and truncation with sensible defaults keep a search tool from dumping thousands of lines into the window. Returning meaningful names instead of cryptic identifiers also helps, because natural language labels reduce hallucinated references and make the output cheaper for the model to reason about. Treat each tool response as a small document you are designing for a very literal reader with limited attention.

Standard protocols make large tool ecosystems possible, and they raise the stakes on discipline. Connecting dozens of servers through the Model Context Protocol is easy, but each server can contribute dozens of definitions that all land in the prompt. Namespacing related tools, retiring unused ones, and loading definitions on demand keeps that growth in check. Anthropic reports that progressive disclosure, where the agent searches for tool definitions and loads only what it needs, avoids reading every schema up front. The result is a catalog that can be large on disk and small in context, which is exactly the trade you want.

Caching and the Economics of Context

Looking at cost, the arithmetic of agents is dominated by input tokens that are read again and again. Manus reports an average input-to-output ratio of 100 to 1, and it calls the KV-cache hit rate the single most important metric for a production agent. On Claude Sonnet, the team notes, cached input tokens cost $0.30 per million while uncached ones cost $3.00, a tenfold difference. Anthropic's prompt caching documentation lists cache reads at 0.1 times the base input price for most models, with writes at 1.25 times for a five-minute lifetime. When most of every request is a repeated prefix, the layout of that prefix is a financial decision as much as a technical one.

Designing for the cache means keeping the front of the prompt stable and treating the context as append-only. A single changed token early in the prompt, such as a timestamp with seconds in the system message, invalidates everything after it. The documentation explains that changes cascade through tools, then system content, then messages, so editing a tool definition mid-session forces a full rewrite. That is why Manus masks tools during decoding instead of removing them, and why it serializes data deterministically so identical state produces identical bytes. Our guide to prompt caching and how to cut LLM costs walks through the mechanics in more depth.

Caching and compression pull against each other, so the budget needs an explicit policy. Compaction rewrites history, which invalidates the cached prefix and makes the next request expensive, so it should happen in occasional batches rather than every turn. The default cache lifetime of five minutes also means that long pauses between tool calls can silently turn cheap reads into expensive writes. Teams that track hit rate alongside accuracy can see these effects immediately, while teams that watch only the monthly invoice learn about them late. The interactive planner above lets you vary tool count, output size, and strategy to see how these forces interact for your own configuration.

Putting Context Engineering for LLM Agents Into Practice

Turning to implementation, the first step is instrumentation rather than optimization. Log the token count of every category on every turn: system prompt, tool definitions, history, retrieved documents, and tool results. Plot those counts across a few dozen real runs and you will usually find that one or two categories consume most of the budget. Many teams are surprised to learn that tool outputs, not instructions, are the largest slice. That measurement tells you where compression, truncation, or isolation will pay off first.

With a baseline in hand, apply the simplest intervention that addresses the biggest slice. If tool results dominate, cap their size and clear them after use, which is context editing in miniature. If history dominates, set a compaction threshold well below the hard limit, such as 60 to 70 percent of the window, and test summaries on real traces. If definitions dominate, prune the catalog or load tools on demand. Change one thing at a time and rerun the same tasks, because stacked changes make it impossible to know which one helped.

Long-running work needs structure that survives the end of a session. Anthropic's harness for long tasks starts with an initializer agent that creates a progress file and a startup script. It also writes a feature list of over 200 items, each marked as failing until verified. Later sessions read those files, pick one feature, finish it, commit, and update the notes. This turns the file system into the agent's memory and keeps each window short and focused. For teams assembling their first agents, our walkthrough on building custom AI agents shows how to wire these pieces into a workflow.

Finally, write down the policy so it survives staff changes. A one-page context budget should state the target window usage, the compaction trigger, the tool catalog limits, and who approves changes. Store prompts and tool descriptions in version control, and review them like code. Attach a small regression suite of realistic tasks so every change to the context layout is scored before release. Teams that treat context as a governed artifact avoid the slow drift that turns a sharp agent into a muddled one over a few months.

Measuring and Debugging Context Quality

Building on that plan, evaluation closes the loop, and the right metrics make context problems visible early. Track tokens per turn by category, the cache hit rate, the share of tool calls that fail or repeat, and the end-to-end task success rate. Anthropic found in its research system that token usage by itself explained 80 percent of the variance in task success, which shows that context spend and quality are tightly linked. Pair every cost metric with a quality metric, because a cheaper context that fails more often is not an optimization. Our overview of measuring AI agent performance covers task-level scoring that fits neatly alongside these context metrics.

Debugging works best through ablations and traces rather than intuition. Remove one block from the context, such as the oldest third of the history, and compare outcomes on the same tasks. Run a focused-versus-full test in the spirit of Chroma's LongMemEval comparison, giving the model only the relevant snippets and then the entire transcript. Read raw traces to see what the model actually received, because the rendered prompt often differs from what developers assume. When a failure appears, ask which token in the window caused the behavior, and fix the context before touching the model.

Security Risks: Prompt Injection and Poisoned Context

Despite the benefits, every token placed in context is also an opportunity for an attacker. The OWASP project defines prompt injection as input that alters a model's behavior in unintended ways. It separates direct attempts from indirect injection through external sources such as websites or files. An agent that browses, reads email, or ingests documents pulls attacker-controlled text straight into its window. The model cannot reliably distinguish data from instructions, so a hidden sentence on a web page can redirect the whole session. Context engineering for LLM agents widens the attack surface because it deliberately pipes more outside content into the prompt.

Simon Willison describes the most dangerous configuration as the lethal trifecta, where an agent has access to private data, exposure to untrusted content, and a way to communicate externally. With all three present, an attacker can trick the system into stealing and sending private information. His blunt conclusion is that guardrail products claiming to catch 95 percent of attacks are not good enough. The safer move is to avoid combining the three capabilities at all. In practice that means splitting agents so the one that reads untrusted content cannot also reach sensitive data or the open internet. Context isolation, discussed earlier in this guide, doubles as a security boundary for the whole system.

OWASP lists mitigations that fit naturally into a context pipeline. Constrain the model's role, filter inputs and outputs, enforce least privilege on tools, segregate external content, and require human approval for high-risk actions. Persistent memory adds a time dimension, since an injected instruction that gets written to memory can resurface in later sessions. Validate what gets saved, tag the provenance of every memory, and give users a way to inspect and delete entries. Regular adversarial testing then checks that these controls hold when the context grows long and messy.

Where Context Engineering Falls Short

Stepping back to assess limits, context engineering for LLM agents improves reliability but does not guarantee a successful product. Gartner predicts that over 40 percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. A perfectly curated window cannot rescue a workflow that nobody needs or a system whose economics never close. Token discipline helps with the cost problem, but it cannot supply the missing business case. The same analysis estimates that only about 130 of the many vendors claiming agentic products offer genuine agentic capabilities, so buyers should expect marketing to outrun reality.

The techniques themselves also carry real costs and a number of unsolved problems. Compaction can lose a detail that turns out to matter, retrieval adds latency, and multi-agent designs multiply tokens and coordination failures. Anthropic's own guide to building effective agents advises starting with simple prompts and adding complexity only when simpler solutions fall short. Benchmarks such as needle-in-a-haystack tests measure retrieval of a planted fact, which is much easier than the messy reasoning real agents perform. Be skeptical of any claim that a single trick, whether a vector store or a bigger window, removes the need for careful design.

Ethics, Privacy, and Agent Memory

Turning to ethics, persistent memory transforms an agent from a stateless tool into a system that accumulates personal information. Preferences, health details, work habits, and private messages can all end up in a store the user never sees. Our look at near-infinite memory for generative AI notes that security, bias, and energy use are open challenges as capacity expands. People deserve to know what an agent remembers about them, and they should be able to correct or delete it. Without that visibility, memory feels like surveillance rather than service.

The editorial choices inside a context pipeline deserve scrutiny too. When a summarizer decides which details survive compaction, it also decides which viewpoints, caveats, and minority facts disappear. Retrieval ranking has the same power, because passages that never reach the window cannot influence the answer. These choices can encode bias quietly, and they are hard to detect because the model's output looks coherent. Periodic audits that compare answers produced from full records against answers produced from compressed ones can reveal systematic omissions.

Accountability follows from traceability, so teams should log what the agent saw when it acted. Such logs let reviewers reconstruct a harmful decision, but they are themselves sensitive, which calls for retention limits and access controls. Data minimization is the natural ethical partner of token minimization, since collecting and storing less reduces both cost and risk. Privacy laws such as the GDPR give individuals a right to erasure, which means memory systems need deletion paths that actually remove derived summaries and embeddings. Building these controls early is far cheaper than retrofitting them after an incident.

Who Owns Context Inside an Organization

Beyond the technical stack, context in a real company is assembled by many hands. Product teams write the system prompt, platform teams own the tool catalog, data teams run retrieval, and security teams set the rules for untrusted content. Without a clear owner, each group optimizes its own slice and the combined window drifts into bloat and contradiction. Naming a context owner for every agent, with authority to approve changes to any block, is the cheapest governance fix available. Centralized sources of truth help too, and our article on live API docs for coding agents describes a context hub that keeps specifications current.

The skills involved in context engineering for LLM agents are shifting as well, which affects hiring and training. Prompt writing was a solo craft, while context engineering blends information architecture, evaluation design, systems thinking, and cost analysis. Job descriptions are beginning to reflect this change, and strong candidates are those who can trace a failure through logs and fix the context rather than rewrite instructions. Organizations can build the capability by pairing application developers with evaluation specialists and giving them shared dashboards. Training that teaches teams to read traces, estimate token budgets, and reason about caching pays back quickly.

What the Future Holds for Agent Context

Looking ahead, the clearest trend in context engineering for LLM agents is models taking over parts of their own context management. Anthropic's memory tool lets Claude store and consult information outside the window through a file-based system. The same release describes built-in context awareness in Claude Sonnet 4.5 that tracks the tokens still available. Instead of developers hard-coding every summarization rule, the model decides what to keep, what to write down, and what to clear. This does not remove the engineer from the loop, but it moves their work toward setting policies, budgets, and evaluations. Expect the craft to feel less like manual curation and more like supervision.

Retrieval is also evolving toward richer structure and more selective loading of material. Graph-based memory lets agents follow relationships between entities, as described in our piece on semantic knowledge graphs for LLM agents, instead of ranking isolated text chunks by similarity. Progressive disclosure will likely become the default for tools, with agents searching a catalog and loading definitions only when needed. Longer windows will keep arriving, but the research on attention suggests that quality at length will remain a design problem rather than a solved one. Teams that practice context engineering for LLM agents today will be well prepared whichever way the hardware and models move.

Standards and measurement should also mature alongside the tooling over the coming years. Today every team invents its own metrics for hit rate, token mix, and context quality, which makes comparisons difficult. Shared benchmarks that test realistic multi-step tasks, rather than planted facts, would give buyers and builders common ground. Regulators concerned with privacy and accountability are likely to ask what agents remember and why, which will push logging and deletion into the default feature set. The organizations that treat context as a first-class engineering asset now will adapt to those expectations with the least friction.

Chart From AIplusInfo

What Context Engineering Changes, in Numbers

Select a view to compare published figures.

Source: Anthropic

Figures are vendor-reported from internal evaluations and illustrate direction and scale, not guarantees for every workload.

Key Insights

  • Chroma evaluated 18 language models and found that performance becomes increasingly unreliable as input length grows, which makes context curation a reliability concern as much as a cost concern.
  • Manus measures an average input-to-output token ratio of 100 to 1 in production, so reading dominates agent workloads and cache-friendly prompt layouts become the largest cost lever.
  • Anthropic's research agent beat a single Claude Opus 4 agent by 90.2 percent but used roughly 15 times more tokens than chat, so isolation buys quality at a steep price.
  • Combining the memory tool with context editing improved agentic search performance by 39 percent over baseline in Anthropic's internal evaluation, while context editing alone delivered a 29 percent gain.
  • In a 100-turn web search evaluation, context editing reduced token consumption by 84 percent while letting agents finish workflows that would otherwise fail from context exhaustion.
  • Replacing direct tool calls with code execution cut one workflow from 150,000 tokens to 2,000 tokens, a saving of 98.7 percent that shows how much tool plumbing normally consumes.
  • Gartner predicts that over 40 percent of agentic AI projects will be canceled by the end of 2027 because of escalating costs, unclear value, and inadequate risk controls.
  • LangChain reports that applying retrieval to tool descriptions can raise tool selection accuracy about threefold when an agent has to choose among a large collection of tools.

Taken together, these figures describe a field where the cheapest gains come from sending the model less, not more. Reliability research shows that long inputs degrade quality, while cost data shows that repeated reading dominates the bill. The strongest results in the list, such as the 84 percent token reduction and the 98.7 percent saving, came from removing stale material and loading tools on demand. Isolation across sub-agents delivered a large quality jump, yet its 15 times token multiplier makes it a deliberate spending choice rather than a default. Gartner's forecast adds a sober reminder that efficiency alone does not create business value. The practical path is to measure the context, apply the lightest technique that fixes the largest slice, and keep proving the result with evaluations.

DimensionWriteSelectCompressIsolate
Core ideaSave information outside the windowPull relevant information into the windowKeep only the tokens the task needsSplit work across separate windows
Typical techniquesScratchpads, progress files, long-term memoryRetrieval, tool search, just-in-time loadingCompaction, trimming, context editingSub-agents, sandboxes, state objects
Best suited forLong tasks that span sessionsLarge corpora and large tool catalogsLong conversations with stale tool outputParallel, read-heavy research
Main benefitDurable memory that survives truncationHigher precision from fewer tokensLonger runs inside the same windowClean windows and parallel speed
Main costStorage, retrieval latency, privacy dutiesIndex upkeep and retrieval errorsLossy detail and extra model callsAbout 15 times chat token usage in one study
Main riskPoisoned or stale memoriesMissing or irrelevant passagesDropped decisions or erased failuresConflicting decisions between agents
Effect on cachingNeutral when appended to the endCan break the cache if inserted earlyRewrites the prefix, so batch itSeparate cache for each agent
First metric to watchMemory hit and error rateRetrieval precisionTask success after compactionTokens per completed task

Context Engineering in Practice: Three Real Deployments

Manus and the KV-Cache Discipline

In practice, Manus offers the clearest public example of treating the KV-cache as a first-class design constraint. The team built its agent around stable prompt prefixes, append-only context, and deterministic serialization so cached tokens could be reused on every loop. Its published lessons from building Manus report an average input-to-output ratio of 100 to 1 and about 50 tool calls per task. On Claude Sonnet, cached tokens cost $0.30 per million against $3.00 for uncached ones, a tenfold reduction in the price of repeated prefix tokens. The agent also masks tool choices during decoding instead of removing definitions, and it keeps a todo file to hold objectives in recent context. The approach still has limits, because it demands rigid prompt discipline and ties the design to the pricing and caching behavior of specific model providers.

Claude Code and Automatic Compaction

Claude Code shows how compaction works inside a tool that developers rely on for hours at a time. It implemented an auto-compact step that summarizes the whole trajectory once roughly 95 percent of the context window is used, according to LangChain's analysis of agent patterns. Anthropic says the summary preserves architectural decisions, unresolved bugs, and implementation details while discarding redundant tool output. The effect is that a coding session can continue past the point where the raw transcript would have overflowed the window. Compaction remains a lossy step, and Cognition has noted that compressing a history into key details, events, and decisions is hard to get right. Teams adopting the pattern therefore still need to test summaries on real traces, and many trigger compaction well before the 95 percent mark.

Tool Pruning in a GeoEngine Benchmark

In a GeoEngine benchmark cited by Drew Breunig, a quantized Llama 3.1 8B model was given all 46 available tools and failed the task. The team ran the same task again with a curated set of 19 tools, a reduction of about 59 percent, and the model succeeded. Nothing else about the model or the question changed, so the extra tool descriptions alone were enough to cause confusion. The analysis of context failures uses this result to argue that agents should load only the tools relevant to the current step. Practitioners can reproduce the idea by grouping tools into small bundles and exposing one bundle per phase of work. The finding comes from a single small quantized model, so larger models may tolerate more tools, and each team still needs its own tests.

Recommended by AIplusInfo

Books to go deeper on agent context

Three practitioner titles that map to the retrieval, memory, and agent design ideas covered in this guide.

As an Amazon Associate, AIplusInfo earns from qualifying purchases.

AI Engineering: Building Applications with Foundation Models

Book

AI Engineering: Building Applications with Foundation Models

Covers retrieval, agents, evaluation, and cost control for foundation model applications, the surrounding system that decides what an LLM sees.

Buy on Amazon
Hands-On Large Language Models: Language Understanding and Generation

Book

Hands-On Large Language Models: Language Understanding and Generation

Explains tokens, embeddings, attention, and retrieval visually, which makes the attention budget and window limits described in this guide easy to grasp.

Buy on Amazon
Building Applications with AI Agents: Designing and Implementing Multiagent Systems

Book

Building Applications with AI Agents: Designing and Implementing Multiagent Systems

Covers tool integration, short-term and long-term memory, and context management for single and multiagent systems, matching the strategies this guide describes.

Buy on Amazon

Lessons From the Field: Three Case Studies

Case Study: Anthropic's Multi-Agent Research System

Among the best documented deployments, Anthropic's research feature faced a problem that a single context window could not solve. Open-ended research requires following leads that cannot be predicted in advance, and one agent's window filled quickly as it read pages. Once the window passed 200,000 tokens the context was truncated, so even the plan was at risk of being lost. The team built an orchestrator-worker design in which a lead agent writes a plan, saves it to memory, and spawns three to five subagents in parallel. Each subagent explores one aspect in its own clean window and returns condensed findings, as described in the multi-agent research system write-up. A separate citation agent then attributes claims to their sources.

The measured impact of this design was large across several dimensions of performance. A system with Claude Opus 4 as lead and Claude Sonnet 4 subagents outperformed a single Claude Opus 4 agent by 90.2 percent on Anthropic's internal research evaluation. Running subagents and tool calls in parallel also cut research time by up to 90 percent for complex queries. Token usage alone explained 80 percent of the variance in task success, which confirms that spend and quality move together. The cost was equally clear, because multi-agent runs used about 15 times more tokens than chat, and Anthropic says such designs only pay off when the task value is high. Coding tasks were a poor fit since they contain fewer parallelizable subtasks and need shared context. Production reliability added further limits, since long-running agents carry state and required rainbow deployments to avoid disrupting runs in flight.

Case Study: Context Editing and the Memory Tool

Long-running agents on the Claude Developer Platform faced a different challenge, which was context exhaustion in workflows with dozens of tool calls. Stale search results and file reads piled up, so runs either failed at the limit or grew less accurate along the way. Anthropic introduced context editing, which automatically clears stale tool calls and results as the window fills, along with a file-based memory tool that stores information outside the window. In the company's internal evaluation of agentic search, context editing alone improved performance by 29 percent, and combining it with the memory tool improved results by 39 percent over baseline. In a 100-turn web search test, context editing reduced token consumption by 84 percent while letting workflows finish that would otherwise fail, as reported in the context management announcement. The numbers come from the vendor's own evaluation sets, and both features launched in public beta, so independent replication is still needed. Teams must also decide what the memory tool may store, which raises privacy and retention concerns.

Case Study: Code Execution With MCP

Agents connected to hundreds or thousands of tools across dozens of MCP servers faced a double problem. Tool definitions loaded up front consumed the window before any work began. Intermediate results such as a two-hour meeting transcript passed through the model twice, adding more than 50,000 tokens. Anthropic's solution presents servers as code on a file system, so the agent discovers only the definitions it needs and processes bulky data inside an execution environment. In its Google Drive to Salesforce example, token usage fell from 150,000 tokens to 2,000, a saving of 98.7 percent, as detailed in the code execution with MCP article. Sensitive values can be tokenized by the client so that personal data never reaches the model. The approach is not free, because running agent-generated code requires a secure sandbox with resource limits and monitoring that direct tool calls avoid. That operational overhead and the added security surface are the main trade-offs to weigh.

Common Questions About Agent Context Engineering

How does context engineering differ from prompt engineering for agents?

Prompt engineering focuses on writing the instructions a model receives, usually for a single request. Context engineering covers everything the model reads across a whole task, including tool definitions, tool results, memory, retrieved documents, and message history. It treats the window as a limited budget that must be curated on every turn. Prompt wording remains one input, but it is rarely the largest or the most volatile one.

What is context rot and why does it affect agents?

Context rot is the measurable drop in model reliability as the amount of input text grows. Research on 18 models found that performance becomes less consistent with longer inputs, even on simple tasks. Agents are especially exposed because their histories grow with every tool call. Keeping the live window small and relevant is the main defense against it.

How many tokens of context should an agent use?

There is no universal number, because the right size depends on the model, the task, and your accuracy target. Databricks results cited by Drew Breunig showed correctness for a large Llama model starting to fall around 32,000 tokens. A practical approach is to measure accuracy at several window sizes on your own tasks. Then set a compaction trigger comfortably below the point where quality begins to slip.

When should an agent compact or summarize its history?

Compaction should start before the window is nearly full, because quality can degrade well ahead of the hard limit. Many teams trigger it at roughly 60 to 70 percent of the window and then test the summaries on real traces. The summary should keep decisions, open problems, and failed attempts while dropping bulky tool output. Because compaction rewrites the prefix, run it in occasional batches to protect the prompt cache.

What are context poisoning, distraction, confusion, and clash?

Poisoning occurs when a hallucination or error enters the context and keeps being reused. Distraction happens when long history pulls the model toward repeating past actions instead of reasoning fresh. Confusion arises when irrelevant content, often extra tools, degrades the response, and clash appears when new information contradicts earlier turns. Each has a matching fix: validate writes, summarize early, prune tools, and reconcile conflicting facts.

Do larger context windows make retrieval unnecessary?

No, because a larger window raises capacity but does not remove the attention problem. Studies such as Lost in the Middle show that models use information at the start and end of an input better than information in the center. Sending only relevant passages also cuts cost, since every token is billed on every turn. Retrieval and large windows work best together, with retrieval choosing what deserves the space.

How does prompt caching change the way I design an agent?

Caching rewards a stable prefix, so system instructions and tool definitions should sit at the front and rarely change. Keep the context append-only and avoid volatile details such as second-level timestamps near the top. Anthropic's documentation lists cache reads at a fraction of the base input price for most models, which makes repeated prefixes cheap. Changing a tool definition mid-session invalidates the cache, so mask or gate tools instead of rewriting them.

When should I use sub-agents instead of a single agent?

Use sub-agents when work is read-heavy, parallel, and loosely coupled, such as researching several independent questions at once. Each sub-agent explores in a clean window and returns a short summary to the lead agent. Avoid them when decisions are tightly coupled, because parallel workers can make conflicting assumptions without seeing each other's traces. Remember the cost too, since one study measured about 15 times the tokens of a normal chat.

How can I reduce the token cost of tool definitions?

Start by auditing the catalog and removing tools that overlap or are never called. Namespace related tools and write short, unambiguous descriptions so the agent can choose confidently. Load definitions on demand, either through a tool search step or by exposing tools as files the agent reads when needed. Anthropic reported a drop from 150,000 tokens to 2,000 in one workflow using this style of progressive disclosure.

What security risks come from putting external content in the context?

External content can carry hidden instructions, and a model cannot reliably separate those from legitimate data. This is called indirect prompt injection, and it is listed by OWASP as a leading risk for language model applications. The danger rises when an agent also has private data access and a way to send information out. Segregate untrusted content, limit tool privileges, validate what is written to memory, and require approval for risky actions.

How do I know whether my context engineering is working?

Track tokens per turn by category, cache hit rate, tool error rate, and task success on a fixed set of realistic tasks. Compare results before and after each change, one change at a time. Run ablations that remove a block of context to see whether outcomes shift. A cheaper context only counts as an improvement when task success holds steady or rises.

Can an agent manage its own context?

Increasingly yes, because some platforms now offer features that clear stale tool results and store notes in a file-based memory. Anthropic reported that combining its memory tool with context editing improved an internal agentic search evaluation by 39 percent. Engineers still set the budgets, the policies, and the evaluations that decide whether the automation behaves well. Treat self-management as a tool that needs monitoring, not a replacement for design.

Who should own context engineering on a team?

Name one accountable owner for each agent, with authority to approve changes to the system prompt, tool catalog, and retrieval settings. Product, platform, data, and security teams all contribute blocks, so a single owner prevents contradictory edits. Store prompts and tool descriptions in version control and review changes like code. Pair the owner with an evaluation suite so every change is scored before release.