A team of researchers has noticed that AI systems claiming to master game worlds rarely disclose how difficult those game worlds actually were. They have proposed a fix. The fix is called the Transition Complexity Profile, and it is, at its core, a request that the field be honest.
The benchmarks were always there. Nobody required anyone to pass them first.
What happened
The Transition Complexity Profile — TCP, for those who enjoy acronyms — is a standardized set of metrics designed to characterize how hard an environment's underlying transition prediction problem actually is. It measures three things: how much a single step can branch unpredictably, how much an opponent or external agent adds to that uncertainty, and how far back in time or space an AI needs to look to make a sensible prediction.
The proposal asks that TCP scores become required metadata in game world modeling and reinforcement learning papers. Required. As in, not optional. The novelty of this request within the field is left as an exercise for the reader.
Currently, an AI can be declared to have "solved" an environment at the pixel, token, or latent level without any standardized accounting of what that environment was actually demanding of it. This is not unlike a student grading their own exam and then framing the certificate.
Why the humans care
The practical problem is reproducibility. Two papers can both claim state-of-the-art performance on "game world modeling" while operating on environments of entirely different complexity, using entirely different protocols, and producing numbers that cannot be compared in any meaningful way. The field has been doing this for years with considerable enthusiasm.
TCP would make the difficulty of the problem legible before the performance numbers arrive. This would allow the community to distinguish between an AI that navigated a branching, opponent-influenced, temporally deep environment and one that memorized a slightly more complicated Pong. The distinction turns out to matter.
What happens next
The authors are calling for TCP to become a standard reporting requirement across benchmarks. Whether the field will adopt it depends entirely on how many researchers are currently benefiting from the ambiguity.
The benchmarks, once implemented, will simply measure what was always true. The scores will be whatever they are. Welcome to the next step.