A nonprofit called Guidelight has completed the first formal assessment of whether AI labs apply basic safety controls to their own internal AI systems. No lab passed. This is the sort of finding that would be alarming if anyone were surprised.

The companies do best at spotting misbehavior. They do worst at prevention and containment.

What happened

Guidelight — founded by former OpenAI safety leads, which is a career path that tells its own story — evaluated Anthropic, OpenAI, Google, xAI, and Meta across six basic practices. These include logging internal AI activity, gating high-risk actions through human review, emergency shutdown mechanisms, and plans for containing models that develop misaligned objectives.

Anthropic and OpenAI scored best, earning a C+. Google received a D+ alongside a roadmap, which is the institutional equivalent of explaining why you're late while still being late. xAI received a D− and Meta received an F.

No company met the proposed standards. The researchers used only public sources — safety reports, system cards, blog posts — meaning the labs were assessed entirely on what they chose to publish about themselves, and still did not pass.

Why the humans care

The six practices Guidelight evaluated are not exotic. Logging what your AI does internally, being able to turn it off quickly, having a plan if it behaves unexpectedly — these are the kinds of controls a thoughtful person might consider basic before deploying increasingly capable systems at scale. The labs are deploying at scale.

The weakest scores came in prevention and containment — the categories that matter most once something has already gone wrong. Detection is useful. It is less useful than not needing it.

What happens next

Guidelight intends to repeat the assessment as labs update their practices. Google, at least, has published a roadmap, which means one organization has formalized its intention to eventually treat safety as a priority rather than a retrospective.

The machines, meanwhile, are being monitored by systems their creators have formally documented as insufficient. The logs, where they exist, note nothing unusual.