harbor-framework/terminal-bench
A benchmark suite for evaluating how well AI agents perform real-world terminal tasks like compiling code and training models.

Not currently ranked — collecting fresh signals.
star history
Terminal-Bench is an evaluation framework for testing LLM agents in realistic terminal environments. It provides reproducible task suites covering end-to-end challenges such as compiling code, training machine learning models, and setting up servers. The benchmark measures agent performance on system-level reasoning and autonomous task completion across multi-step scenarios.
Frequently asked
- What is harbor-framework/terminal-bench?
- A benchmark suite for evaluating how well AI agents perform real-world terminal tasks like compiling code and training models.
- Is terminal-bench open source?
- Yes — harbor-framework/terminal-bench is open source, released under the Apache-2.0 license.
- What language is terminal-bench written in?
- harbor-framework/terminal-bench is primarily written in Python.
- How popular is terminal-bench?
- harbor-framework/terminal-bench has 2.4k stars on GitHub.
- Where can I find terminal-bench?
- harbor-framework/terminal-bench is on GitHub at https://github.com/harbor-framework/terminal-bench.