← all repositories
TIGER-AI-Lab/ClawBench

Browser agents still fail two-thirds of everyday web tasks

ClawBench measures whether AI browser agents can handle real-world online tasks—booking flights, ordering food, applying for jobs—on live websites rather than sanitized sandboxes.

505 stars Python LLMOps · EvalAgents
ClawBench
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

What it does

ClawBench is an open-source evaluation framework that pits AI browser agents against 283 everyday tasks across 144 live websites. It covers mundane but complex activities like travel booking, food ordering, and job applications, scoring end-to-end success through a five-layer recording pipeline and an agentic evaluator that checks runs against human references. The benchmark is packaged as a PyPI module and ships with full execution traces on Hugging Face.

The interesting bit

Instead of testing agents on static HTML snapshots or toy forms, ClawBench runs them against real, changing websites with live accounts and actual checkout flows. The scoring pipeline captures everything from network requests to final screenshots, then grades results with both an automated rubric and an inline LLM judge—because determining whether a pizza was actually ordered requires more than a string match.

Key highlights

  • 283 tasks split across V1 (153 tasks) and V2 (130 tasks), spanning 15 life categories
  • Runs on live websites using Docker-isolated harnesses and a request interceptor
  • Top-performing agents complete roughly one in three tasks (33.3% success rate to date)
  • Ships with three Hugging Face datasets including full five-layer execution traces for V1 and V2
  • Supports multiple agent harnesses including browser-use and hermes

Caveats

  • V2 introduces a “lenient judge” and new metrics (reward and intercepted), so scores are not directly comparable to V1’s 33.3% top rate

Verdict

Researchers building browser agents should care because it replaces toy benchmarks with live checkout flows and actual inboxes. If you are not working on web agents or agent evaluation, treat this as a public scoreboard for how far end-to-end automation still has to go.

Frequently asked

What is TIGER-AI-Lab/ClawBench?
ClawBench measures whether AI browser agents can handle real-world online tasks—booking flights, ordering food, applying for jobs—on live websites rather than sanitized sandboxes.
Is ClawBench open source?
Yes — TIGER-AI-Lab/ClawBench is open source, released under the Apache-2.0 license.
What language is ClawBench written in?
TIGER-AI-Lab/ClawBench is primarily written in Python.
How popular is ClawBench?
TIGER-AI-Lab/ClawBench has 505 stars on GitHub.
Where can I find ClawBench?
TIGER-AI-Lab/ClawBench is on GitHub at https://github.com/TIGER-AI-Lab/ClawBench.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.