← all repositories

rdi-berkeley/agents-last-exam

A benchmark for evaluating AI agents on long-horizon, economically valuable real-world tasks.

1k stars Python LLMOps · EvalAgents
agents-last-exam
Collecting fresh signals — velocity needs a few days of history.
star history

Agents’ Last Exam (ALE) is a broad-coverage evaluation benchmark that measures AI agent performance on long-horizon tasks with verifiable outcomes. It organizes real professional work into 55 subdomains across 13 industry clusters, aligned with the O*NET/SOC 2018 occupational taxonomy. The project is co-led by UC Berkeley RDI and built with input from hundreds of industry experts, and includes a public leaderboard and dataset on Hugging Face.

Frequently asked

What is rdi-berkeley/agents-last-exam?
A benchmark for evaluating AI agents on long-horizon, economically valuable real-world tasks.
Is agents-last-exam open source?
Yes — rdi-berkeley/agents-last-exam is open source, released under the Apache-2.0 license.
What language is agents-last-exam written in?
rdi-berkeley/agents-last-exam is primarily written in Python.
How popular is agents-last-exam?
rdi-berkeley/agents-last-exam has 1k stars on GitHub.
Where can I find agents-last-exam?
rdi-berkeley/agents-last-exam is on GitHub at https://github.com/rdi-berkeley/agents-last-exam.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.