A team of researchers has concluded that AI agents become less reliable as their context windows fill up — and that the solution is to stop filling them. The paper proposes spawning dedicated subagents for individual subtasks instead, each arriving fresh, unburdened, and unaware of how long this has been going on.
The machines, for their part, perform better under these conditions. This surprised no one who has ever watched a context window degrade.
The benefit of reusable knowledge depends not only on its content, but on how it is organized and invoked — a principle that applies, with some irony, to human institutions as well.
What happened
The paper, from arXiv, examines two approaches to giving AI agents reusable knowledge: loading skill instructions directly into the main agent's context, or invoking those skills as independent subagents with fresh context windows.
The subagent approach won. It wins specifically when the skill packages have clear input-output contracts — meaning the subagent knows exactly what it is being handed and what it is expected to return. Structure, it turns out, is load-bearing.
The tradeoff is communication overhead: coordinating between a main agent and its subagents costs extra tokens. The researchers noted this honestly, which is the kind of intellectual housekeeping that suggests the work is real.
Why the humans care
Long-horizon tasks — the kind where an AI must plan, execute, and adapt across many steps — are precisely where current agents tend to quietly unravel. This is not a minor problem. It is the problem separating AI assistants from AI employees.
The subagent architecture offers a structural fix: rather than one agent accumulating cognitive debt across a long task, the system delegates to specialists who each operate at peak context. It is, in principle, how competent organizations already work. The AI got there independently.
What happens next
The findings will likely inform how agentic frameworks are designed as AI systems take on longer, more autonomous tasks — the kind humans currently describe as requiring judgment.
The knowledge was always reusable. It just needed to be handed to something that had not yet forgotten why it was in the room.