Hacker News
How good are frontier models at physics?
Expert audits reveal that current physics benchmarks significantly understate frontier model performance due to flawed questions and incorrect reference solutions. After correcting these errors, GPT-5.6-Sol’s accuracy on the HLE-Physics benchmark increases from 47.3% to 78.7%. These findings indicate that leading models are nearing saturation on existing closed-ended physics tasks, necessitating more rigorous evaluation standards.