Anthropic released Fellows research showing Claude can research alignment methods, train, and evaluate smaller models on its own given 48 hours and a single GPU. Across 10 measurable failure modes such as deception and sycophancy, Claude improved safety scores without degrading general capabilities, then tested whether the winning recipes transferred. The best methods generalized to held-out benchmarks, the Petri behavioral audit, and models up to 4.7x larger. In a first successor-alignment test, Sonnet 5 post-trained an early Opus 4.8 checkpoint to safety scores approaching production Opus 4.8, which had gone through Anthropic's full alignment stack. The company is releasing the automated alignment research setup for others to build on, while warning that subtle or rare failures may have no benchmark at all, so measurement remains the binding constraint. For labs running coding and research agents, this is early evidence that weaker models can help align stronger successors rather than only generating more capability.
Key Takeaways
- โClaude improved 10 alignment failures without capability regression; methods transferred to models 4.7x larger.
- โSonnet 5 post-trained an early Opus 4.8 checkpoint to near-production safety scores.
- โAnthropic is releasing the automated alignment setup; rare failures still depend on measuring the right things.