HomePeopleCompaniesAI ModelsOpen SourceAgentsResearchApps
AllPapersBenchmarks
ResearchBenchmarks 28 Aug 2026 HN

Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

TLDR

Terminal-Bench-Science is a new benchmark for evaluating AI agents on scientific research workflows executed entirely in a terminal environment. Tasks include literature search, data analysis, hypothesis generation, and experimental design, assessed against expert researcher baselines. The benchmark is designed to measure autonomous research capability rather than isolated question-answering performance.

Read the original