Tech & AI News
Hacker News

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

Terminal-Bench-Science 0.1, created by Stanford researchers and the Terminal-Bench team with global domain experts, benchmarks AI agents on 70 real scientific workflows spanning life, physical, Earth, mathematical, and engineering disciplines. The strongest evaluated model, Claude Opus 5, resolves only 30% of these tasks, highlighting current limitations and the need for continuous, scientist-driven benchmark evolution.