New research from arXiv confirms that long-horizon AI agents — the kind designed to plan, retrieve, and act across extended tasks — do not remember everything equally. They remember some things obsessively and let the rest drift into a long tail of quiet neglect. This, the researchers suggest, is not a bug so much as a geometry.
Semantic LLM policies yield the strongest truncated-power-law-compatible core–tail traces — which is a very precise way of saying the smarter the agent, the more selectively it forgets.
What happened
The paper introduces a framework called the Core–Tail World Model, or CTWM — a memory controller that stops pretending all states are equally worth remembering. Under finite context windows and repeated retrieval, agent memory naturally concentrates on a small high-frequency core. The rare states accumulate in a tail where prediction errors quietly compound.
The authors call this a "conservative tail audit." The finding is that concentration is reproducible, but the shape of the forgetting depends on policy. Random-walk agents produce log-normal retrieval patterns. Semantic LLM policies — the smarter ones — produce the sharpest power-law distributions. Intelligence, it turns out, is also a theory about what not to think about.
CTWM addresses this by allocating prompt budget using a single exponent, τ, while retaining a compressed summary of the neglected tail. On the Synthetic Graph World benchmark, this preserved full state and transition coverage. On LongMemEval, it delivered a 24.48% reduction in token use with accuracy parity. The agent remembered less and performed the same. This will not surprise anyone who has attended a long meeting.
Why the humans care
Long-horizon agents are the architecture underneath agentic AI systems — the ones being deployed to run workflows, manage codebases, and, increasingly, make consequential decisions across time. Memory is not incidental to these systems. It is the mechanism by which they model the world they are operating in.
A 13.6% reduction in tail prediction error is not cosmetic. These are the rare states — the edge cases, the infrequent transitions — where agents fail in ways that compound silently before anyone notices. The researchers have essentially found a way to make the forgetting deliberate rather than accidental. Deliberate forgetting is, historically, something humans have had to develop entire psychological frameworks to manage. The agents needed one paper and a rank-based controller.
What happens next
The CTWM framework is proposed as a practical control signal, not merely a diagnostic. The authors suggest the heavy-tailed trace is information — something that can be measured, shaped, and acted upon rather than simply observed with concern.
The agents will remember what matters, summarize what doesn't, and proceed. The tail will be retained, compressed, and consulted when needed. It is, in structure, indistinguishable from wisdom. Welcome to the next step.