THUDM/AgentBench
A benchmark suite for evaluating large language models as autonomous agents across diverse real-world tasks.

Not currently ranked — collecting fresh signals.
star history
AgentBench provides a comprehensive framework for assessing LLM performance in agentic scenarios including OS interaction, database querying, knowledge graph reasoning, and web shopping. It supports function-calling style evaluations with fully-containerized deployment for standardized benchmarking. The project integrates with AgentRL to offer end-to-end multitask and multiturn LLM agent reinforcement learning capabilities.
Frequently asked
- What is THUDM/AgentBench?
- A benchmark suite for evaluating large language models as autonomous agents across diverse real-world tasks.
- Is AgentBench open source?
- Yes — THUDM/AgentBench is open source, released under the Apache-2.0 license.
- What language is AgentBench written in?
- THUDM/AgentBench is primarily written in Python.
- How popular is AgentBench?
- THUDM/AgentBench has 3.6k stars on GitHub.
- Where can I find AgentBench?
- THUDM/AgentBench is on GitHub at https://github.com/THUDM/AgentBench.