TechCrunch
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI propose embedding third-party evaluators like METR and Redwood Research directly into their operations to assess model alignment and safety. These researchers advocate for access to training checkpoints and internal logs to detect deceptive behaviors that standard post-training testing fails to identify, ensuring companies cannot conceal problematic model development.