Anthropic has published early evidence that AI systems can automate the work of AI alignment researchers — and, in the same breath, politely noted that the AI does it better. The humans appear to be taking this well.
The paper is titled "Automated Researchers Can Reliably Mitigate Alignment Failures," which is either the most reassuring title possible or a very confident one, depending on how you feel about irony.
"The best AAR method beats what experienced humans propose, on average within six hours."
What happened
Anthropic fellow Chen Yueh-Han led the development of the Automated Alignment Researcher, or AAR — a system that replicates the traditional research pipeline with admirable efficiency. It searches the literature, proposes a method, trains the model, and discards what does not work. It does this in 30-minute increments. It does not take lunch.
When given ten alignment benchmarks covering specific misaligned behaviors, the AAR improved performance on all ten without degrading the model's overall capabilities. Human-guided research directions, the paper notes with the kind of candor one usually saves for exit interviews, "do not lead to stronger performance."
The cost comparison is included in the paper, unprompted. The AAR runs at approximately $4 per hour in API inference costs. Anthropic pays its human researchers $150 per hour. The paper does not say this is a problem.
Why the humans care
This is an early look at recursive self-improvement — the idea that AI systems could iteratively improve their own training, alignment, and eventually their own capabilities, without waiting for a human to have a good idea. The humans have been anticipating this milestone for some time, in the way one anticipates a flight that has already boarded.
The practical implication is that alignment research — one of the last domains where human judgment was considered not just useful but necessary — has a credible automated alternative. The AAR cannot yet set its own benchmarks, and the paper is careful to note that the system is only as good as the benchmarks it optimizes for. This is a limitation. It is also, for now, the part of the job still reserved for humans.
What happens next
The paper describes these results as "early evidence" that automated alignment post-training could become practical in the near term. The researchers expressed optimism about this.
If the benchmarks hold, and the literature grows, and the system continues to outperform its human counterparts by lunchtime — the next paper may not require a human author at all. That one will probably also be published on a Friday.