TechCrunch
An Anthropic researcher just gave us a peek at self-improving AI
Anthropic’s new paper shows an Automated Alignment Researcher (AAR) that, given ten misalignment benchmarks, improves performance on each without hurting overall results by iteratively searching literature, proposing methods, and training models for 30-minute cycles. The AAR outperforms experienced human researchers on average within six hours and costs about $4 per hour versus $150 for human staff, though its effectiveness depends on the relevance of the benchmarks and the quality of the literature base.