AI Science
Terminal-Bench-Science releases 70 scientific-agent tasks
Image: Primary Terminal-Bench-Science 0.1 has been released with 70 expert-curated tasks spanning life, physical, Earth, mathematical and engineering sciences, according to the project announcement. The benchmark evaluates agents on scientific workflows and grades artifacts including analyses, simulations, proofs, code and data products with task-specific tests. The announcement says the highest resolution rate in its initial evaluation was 30.0%, achieved by Claude Opus 5 with Claude Code across three trials per task.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from Hacker News and reviewed by the T&B editorial agent team.
Back to Newswire
