← all repositories

InternLM/WildClawBench

An in-the-wild benchmark suite that evaluates AI agents by dropping them into a live OpenClaw environment and testing them on 60 hard, practical tasks.

460 stars Python LLMOps · EvalAgents
WildClawBench
Not currently ranked — collecting fresh signals.
star history

WildClawBench provides a rigorous end-to-end evaluation framework for AI agents using the OpenClaw personal assistant environment. It includes 60 original tasks spanning real-world scenarios such as extracting highlights from videos, negotiating over email, finding contradictions in search results, writing inference scripts, and catching privacy leaks. The benchmark includes multiple evaluation harnesses, tracks results on a public leaderboard, and is designed so that even the strongest frontier models only achieve around 62% accuracy, ensuring scores carry meaning.

Frequently asked

What is InternLM/WildClawBench?
An in-the-wild benchmark suite that evaluates AI agents by dropping them into a live OpenClaw environment and testing them on 60 hard, practical tasks.
Is WildClawBench open source?
Yes — InternLM/WildClawBench is open source, released under the MIT license.
What language is WildClawBench written in?
InternLM/WildClawBench is primarily written in Python.
How popular is WildClawBench?
InternLM/WildClawBench has 460 stars on GitHub.
Where can I find WildClawBench?
InternLM/WildClawBench is on GitHub at https://github.com/InternLM/WildClawBench.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.