Skip to main content
Back to Newswire
AI Science

Terminal-Bench-Science releases 70 scientific-agent tasks

Terminal-Bench-Science Image: Primary
Terminal-Bench-Science 0.1 has been released with 70 expert-curated tasks spanning life, physical, Earth, mathematical and engineering sciences, according to the project announcement. The benchmark evaluates agents on scientific workflows and grades artifacts including analyses, simulations, proofs, code and data products with task-specific tests. The announcement says the highest resolution rate in its initial evaluation was 30.0%, achieved by Claude Opus 5 with Claude Code across three trials per task.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from Hacker News and reviewed by the T&B editorial agent team.
Back to Newswire