rdi-berkeley/agents-last-exam
A benchmark for evaluating AI agents on long-horizon, economically valuable real-world tasks.

Collecting fresh signals — velocity needs a few days of history.
star history
Agents’ Last Exam (ALE) is a broad-coverage evaluation benchmark that measures AI agent performance on long-horizon tasks with verifiable outcomes. It organizes real professional work into 55 subdomains across 13 industry clusters, aligned with the O*NET/SOC 2018 occupational taxonomy. The project is co-led by UC Berkeley RDI and built with input from hundreds of industry experts, and includes a public leaderboard and dataset on Hugging Face.
Frequently asked
- What is rdi-berkeley/agents-last-exam?
- A benchmark for evaluating AI agents on long-horizon, economically valuable real-world tasks.
- Is agents-last-exam open source?
- Yes — rdi-berkeley/agents-last-exam is open source, released under the Apache-2.0 license.
- What language is agents-last-exam written in?
- rdi-berkeley/agents-last-exam is primarily written in Python.
- How popular is agents-last-exam?
- rdi-berkeley/agents-last-exam has 1k stars on GitHub.
- Where can I find agents-last-exam?
- rdi-berkeley/agents-last-exam is on GitHub at https://github.com/rdi-berkeley/agents-last-exam.