← all repositories
harbor-framework/terminal-bench-science

Scientists are writing the exams they want AI to pass

A benchmark where working scientists contribute the terminal-based research workflows they want AI agents to master.

terminal-bench-science
Collecting fresh signals — velocity needs a few days of history.
star history

What it does

Terminal-Bench-Science evaluates AI agents on 70 expert-curated research tasks spanning life sciences, physics, earth science, mathematics, and engineering. Each task runs in a terminal environment inside Docker and must produce an objectively verifiable result. The project functions as a living benchmark: working scientists contribute new tasks, which pass through automated QA and human review before landing on a public leaderboard.

The interesting bit

Most AI benchmarks feel synthetic; this one is crowdsourced from researchers who want AI to support their actual work. The setup enforces a feedback loop between scientific needs and agent development, with tasks that are genuinely challenging for frontier models. The review pipeline is unusually rigorous—automated checks include a 39-criteria rubric, adversarial “cheat” trials to catch reward hacking, and oracle/no-op validation to ensure tasks are solvable but not trivial.

Key highlights

  • 70 tasks across five scientific domains, growing toward 100+
  • Every task is authored by domain experts and verified in a sandboxed terminal environment
  • Automated QA pipeline: static checks, similarity detection, Docker builds, multi-agent trials, and adversarial validation
  • Continuous benchmark model designed to evolve alongside frontier AI capabilities
  • Public leaderboard and task dashboard for tracking proposals, reviews, and agent performance

Caveats

  • Requires the Harbor framework and a sandboxing provider like Modal or Daytona to run
  • Contribution bar is high: most tasks need several rounds of review and revision before merge
  • Currently at 70 tasks, so coverage is broad but still expanding

Verdict

Worth tracking if you build or evaluate scientific AI agents, or if you are a researcher who wants AI tools tested against real workflows. Skip it if you are looking for a static, off-the-shelf dataset—this is infrastructure with a maintenance overhead.

Frequently asked

What is harbor-framework/terminal-bench-science?
A benchmark where working scientists contribute the terminal-based research workflows they want AI agents to master.
Is terminal-bench-science open source?
Yes — harbor-framework/terminal-bench-science is open source, released under the Apache-2.0 license.
What language is terminal-bench-science written in?
harbor-framework/terminal-bench-science is primarily written in Python.
How popular is terminal-bench-science?
harbor-framework/terminal-bench-science has 565 stars on GitHub.
Where can I find terminal-bench-science?
harbor-framework/terminal-bench-science is on GitHub at https://github.com/harbor-framework/terminal-bench-science.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.