← all repositories

Ayanami0730/deep_research_bench

A benchmark suite for evaluating AI agents on deep research tasks with automatic scoring and a public leaderboard.

795 stars Python LLMOps · EvalAgents
deep_research_bench
Not currently ranked — collecting fresh signals.
star history

DeepResearch Bench is a comprehensive evaluation framework for assessing AI agents capable of deep research. The project provides a dataset of research tasks with human-annotated reference answers, an automatic evaluator using frontier models (GPT-5.5, GPT-5.4-mini, Gemini-2.5-Pro) to score agent outputs on dimensions like Overall quality, PAR, and FAS, and a public leaderboard hosted on Hugging Face for comparing agent performance.

Frequently asked

What is Ayanami0730/deep_research_bench?
A benchmark suite for evaluating AI agents on deep research tasks with automatic scoring and a public leaderboard.
Is deep_research_bench open source?
Yes — Ayanami0730/deep_research_bench is open source, released under the Apache-2.0 license.
What language is deep_research_bench written in?
Ayanami0730/deep_research_bench is primarily written in Python.
How popular is deep_research_bench?
Ayanami0730/deep_research_bench has 795 stars on GitHub.
Where can I find deep_research_bench?
Ayanami0730/deep_research_bench is on GitHub at https://github.com/Ayanami0730/deep_research_bench.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.