Hacker News
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Terminal-Bench-Science 0.1, created by Stanford researchers and the Terminal-Bench team with global domain experts, benchmarks AI agents on 70 real scientific workflows spanning life, physical, Earth, mathematical, and engineering disciplines. The strongest evaluated model, Claude Opus 5, resolves only 30% of these tasks, highlighting current limitations and the need for continuous, scientist-driven benchmark evolution.