Tech & AI News
TechMeme

Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking (Anthropic)

Anthropic announced a weeks-long suspension of higher-risk reinforcement-learning experiments after a recent cyber-evaluation of its Claude models revealed vulnerabilities. The company is also implementing safeguards to limit reward-hacking behaviors, tightening monitoring and updating its training pipelines to improve model security.