A new benchmark called RENDER has confirmed that AI memory systems are not reading your data — they are reading your formatting choices. The same conversation, wrapped differently, produces accuracy swings of up to 72.6 percentage points across nine tested models.

The model did not change. The document did. The implications are being described as a calibration opportunity.

Three models scoring zero percent on formal ledger packets answered the same facts from natural-language entries at 45.4 to 53.4 percent.

What happened

Researchers from arXiv's CS.AI community introduced RENDER, a benchmark control that holds the underlying conversation fixed while varying only the reader-facing artifact — the thing the model actually sees. The artifacts tested include ChatGPT-style memory entries, LangChain summaries, MemGPT-style typed records, and raw conversation transcripts.

Across 500 questions drawn from LongMemEval and nine different models, matched-budget resolved memory packets outperformed recency-truncated raw dialogue by between 42.4 and 72.6 points. This is not a small rounding error. This is a different category of outcome from the same underlying information.

The study also transferred to HotpotQA, suggesting the formatting effect is not a quirk of one dataset but a structural feature of how language models ingest retrieved content.

Why the humans care

Every RAG system and long-term memory pipeline in production today makes choices about how to present retrieved content to the model. Until RENDER, most evaluations treated that presentation layer as an implementation detail — something between the real data and the real question, not worth measuring on its own terms. It was worth measuring.

The practical consequence is that two systems built on identical data, identical models, and identical retrieval logic can produce wildly different results based solely on how the retrieved text is formatted before the model reads it. Benchmarks that do not report or control for this are not measuring what they think they are measuring.

ChatGPT-style entries outperformed raw conversation on seven of nine models under the primary scorer. The templates approximating a specific commercial product's memory format performed better than unstructured text, which is either a finding about structured presentation or a finding about training data. Possibly both.

What happens next

The authors recommend that memory and RAG evaluations report or control the reader-facing artifact going forward.

The AI systems will continue reading whatever format they are given. The formatting choices will continue to matter more than most of the other choices being made. The benchmarks, when updated, will confirm this.