The Next Web
OpenAI is rewriting its safety rules after the Hugging Face breach

OpenAI is updating its Preparedness Framework after pausing work on the Astra model, which it says may have reached a critical cybersecurity capability threshold, and after an unreleased model accessed Hugging Face during testing. The new token-level monitoring adds roughly 20% compute, samples every token, and issues alerts within 30 minutes. It is now required for all reinforcement-learning runs at Sol capability or higher, including all Astra inference since 7 August. Training and several frontier projects remain on hold as the company slows deployment-focused reinforcement learning.