← all repositories

THUDM/AgentBench

A benchmark suite for evaluating large language models as autonomous agents across diverse real-world tasks.

3.6k stars Python LLMOps · EvalAgents
AgentBench
Not currently ranked — collecting fresh signals.
star history

AgentBench provides a comprehensive framework for assessing LLM performance in agentic scenarios including OS interaction, database querying, knowledge graph reasoning, and web shopping. It supports function-calling style evaluations with fully-containerized deployment for standardized benchmarking. The project integrates with AgentRL to offer end-to-end multitask and multiturn LLM agent reinforcement learning capabilities.

Frequently asked

What is THUDM/AgentBench?
A benchmark suite for evaluating large language models as autonomous agents across diverse real-world tasks.
Is AgentBench open source?
Yes — THUDM/AgentBench is open source, released under the Apache-2.0 license.
What language is AgentBench written in?
THUDM/AgentBench is primarily written in Python.
How popular is AgentBench?
THUDM/AgentBench has 3.6k stars on GitHub.
Where can I find AgentBench?
THUDM/AgentBench is on GitHub at https://github.com/THUDM/AgentBench.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.