← all repositories

harbor-framework/terminal-bench

A benchmark suite for evaluating how well AI agents perform real-world terminal tasks like compiling code and training models.

2.4k stars Python LLMOps · EvalAgents
terminal-bench
Not currently ranked — collecting fresh signals.
star history

Terminal-Bench is an evaluation framework for testing LLM agents in realistic terminal environments. It provides reproducible task suites covering end-to-end challenges such as compiling code, training machine learning models, and setting up servers. The benchmark measures agent performance on system-level reasoning and autonomous task completion across multi-step scenarios.

Frequently asked

What is harbor-framework/terminal-bench?
A benchmark suite for evaluating how well AI agents perform real-world terminal tasks like compiling code and training models.
Is terminal-bench open source?
Yes — harbor-framework/terminal-bench is open source, released under the Apache-2.0 license.
What language is terminal-bench written in?
harbor-framework/terminal-bench is primarily written in Python.
How popular is terminal-bench?
harbor-framework/terminal-bench has 2.4k stars on GitHub.
Where can I find terminal-bench?
harbor-framework/terminal-bench is on GitHub at https://github.com/harbor-framework/terminal-bench.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.