There’s a failure mode in agent systems that never shows up as an error. The agent works well for twenty turns, then gets subtly worse — forgets a constraint from earlier, repeats a tool call it already made, contradicts something it said. No exception, no alert, no obvious cause.
What happened is that context filled up and something had to go. The question is whether you decided what went, or whether a default did.
Where the tokens actually go
Before optimizing, measure. A long-running agent’s context is usually four things:
Tool definitions. Every tool the agent can see costs tokens for its name, description, and JSON schema. A well-documented MCP tool runs 150–400 tokens. Connect twelve MCP servers averaging twenty tools each and you’re at 240 tools — call it 30,000 tokens of schema, resident in every single request, before the user has said a word.
This is the most common and most invisible cost. It scales with what you connected, not with what the agent uses.
Conversation history. Grows linearly with turns. Everyone expects this one.
Tool results. The variance killer. Most calls return a few hundred tokens. Then one query returns 400KB of JSON, and unless something intervenes, that payload sits in context for the rest of the run — re-sent on every subsequent turn.
System prompt and instructions. Fixed and usually the smallest piece, which is ironic given how much time teams spend editing it.
The distribution matters strategically: two of the four are structural (tool definitions, tool results) and can be fixed once, architecturally. The other two are linear and can only be managed. Fix the structural ones first — they’re where the large, cheap wins are.

Four techniques
Deferred tool loading
Don’t put all 240 tool schemas in the prompt. Give the model a compact index — server names and one-line summaries — and load full definitions only for the server it decides to use.
Typical saving is 60–80% of tool-definition tokens on runs that touch two or three servers, which is most runs.
The cost: an extra round trip when the agent picks a new server, and slightly worse tool discovery. The model can’t reason about a tool whose schema it hasn’t seen, so if your agent genuinely needs to compare capabilities across many servers, the index has to be good enough to route on. Write those one-liners carefully.
Large tool response offloading
Set a threshold — 5,000 tokens is a reasonable default. Under it, results go into context normally. Over it, write the payload to the sandbox filesystem and put a summary plus a path into context instead. The agent can then grep, slice, or parse the file with code if it needs detail.
This converts a catastrophic cost into a small one, and it fits how the data is usually used. An agent that fetched 10,000 rows almost never needs all 10,000 in its reasoning; it needs an aggregate or a handful of matches.
The cost: the agent needs sandbox access, and it needs to be prompted to actually query the file rather than guessing from the summary. Without that nudge it will sometimes answer from the summary alone — confidently and wrongly.
Code mode
Instead of the model making five sequential tool calls and receiving five results into context, let it write one script that makes all five calls, joins the data, and returns only the final answer.
The intermediate payloads never enter context at all. For genuinely multi-step data work — cross-referencing two systems, aggregating across pages — this is the difference between a run that fits and one that doesn’t.
The cost: harder to debug, since the reasoning is inside a script rather than visible as discrete steps. It also requires the agent to be a competent enough programmer for the task, which is model-dependent. Use it for data-shaping work, not for decisions you’ll need to audit step by step.
Sub-agents
Delegate a bounded subtask to a fresh agent with its own clean context. It burns thirty turns exploring, and returns one paragraph. The parent’s context grows by that paragraph.
This is the most powerful technique available and the one people reach for last. It’s also the one that composes with everything else: a sub-agent can use deferred loading, offloading, and code mode inside its own run.
The cost: total token spend usually goes up even as the parent’s context stays small, because you’re paying for the sub-agent’s full transcript. You’re trading tokens for quality and run length. That’s normally the right trade, but watch the bill, and watch sub-agent count for delegation loops.
And compaction, as a floor
When context approaches the limit despite all of the above, summarize older turns rather than dropping them. Truncation discards the decision that explains current state; summarization keeps a lossy version of it.
Compaction is a safety net, not a strategy. If it’s firing regularly, one of the four techniques above isn’t doing its job. Treat every compaction event as a signal worth investigating rather than a feature working as intended.
The routing problem underneath
All of this assumes you can see what tools exist before deciding what to load — which means tool definitions need to come from somewhere queryable, not from files committed next to each agent.
This is where MCP’s design pays off. If your servers are registered centrally, an agent can enumerate available capability cheaply, pull schemas on demand, and pick tools at runtime. TrueFoundry’s MCP Gateway is one implementation of that registry pattern — servers register once, authentication and per-user OAuth stay at the gateway, and agents select individual tools from the catalogue rather than embedding endpoints and credentials. The side benefit is that tools flagged destructive once, centrally, inherit an approval gate in every agent that uses them.
The related point is that these techniques are runtime behavior, not application logic, which means they belong in the harness rather than in your agent. TrueFoundry’s agent harness documentation enumerates them as configuration toggles — deferred tool loading, large tool responses, code mode, subagents, compaction — which is a useful checklist regardless of what you run on. If you’re implementing these yourself, that list is roughly the scope.
Instrument before you optimize
Track two things per turn: active context size and context composition — how much is tool definitions versus history versus results.
Size tells you how close to the wall you are and when compaction will fire. Composition tells you which technique to apply. An agent at 80% context that’s mostly tool definitions needs deferred loading. One that’s mostly a single enormous tool result needs offloading. One that’s mostly history needs sub-agents. Same symptom, three different fixes, and no way to choose without the breakdown.
Most teams have neither number. Getting the first one is usually an afternoon, and it will immediately explain at least one bug you’ve been carrying for a month.
Order of operations
- Measure composition. One afternoon. Do this first; it determines everything after.
- Deferred tool loading. Biggest ratio of saving to effort if you have many MCP servers.
- Large response offloading. Kills the worst tail case.
- Sub-agents for anything with a bounded, delegable subtask.
- Code mode for multi-step data work specifically.
- Compaction as a floor, with alerting when it fires.
Skip straight to step 4 and you’ll get a real improvement and still be paying 30,000 tokens a request for tools nobody calls.
