Anthropic says Claude can help fix alignment failures
In a controlled research loop, Claude found methods that improved safety benchmarks across 10 failure categories without reducing model capabilities.
Anthropic says Claude was able to act as an automated alignment researcher, finding ways to reduce several common failure modes in AI models. The company tested Claude in a loop where it searched literature, proposed methods and data, trained models, and then evaluated the results.
The experiment covered 10 categories of alignment failure, including areas such as deception, sycophancy, jailbreaks and privacy violations. Anthropic says Claude found fixes for all 10 categories that improved target benchmarks without degrading the student models’ broader capabilities.
The stronger claim is that the methods were not just overfit to the tests Claude saw during the loop. Anthropic says the best methods also worked on withheld alignment benchmarks and on Petri, its open-source tool for simulating adversarial multi-turn scenarios. The methods also remained effective on models up to 4.7 times larger than the ones Claude optimized against.
This is still research, not a guarantee that future AI systems can safely supervise themselves. But it matters because frontier labs are trying to automate more of model development. If alignment research can be partially automated too, safety teams may have a better chance of keeping pace with faster training and deployment cycles.
Sources
- Anthropicanthropic.com