← all repositories

eval-sys/mcpmark

A benchmark suite that stress-tests AI agents by executing one-command tasks across real MCP tool environments with isolated sandboxes.

433 stars Python LLMOps · EvalAgents
mcpmark
Not currently ranked — collecting fresh signals.
star history

MCPMark provides a reproducible evaluation framework for measuring AI agent performance in real-world MCP tool use scenarios. The benchmark runs agents against tasks in environments including Notion, GitHub, Filesystem, Postgres, and Playwright, with isolated sandbox execution and automatic recovery for failures. It generates unified metrics and aggregated reports for comparing model capabilities, and supports trajectory logging to Hugging Face datasets.

Frequently asked

What is eval-sys/mcpmark?
A benchmark suite that stress-tests AI agents by executing one-command tasks across real MCP tool environments with isolated sandboxes.
Is mcpmark open source?
Yes — eval-sys/mcpmark is open source, released under the Apache-2.0 license.
What language is mcpmark written in?
eval-sys/mcpmark is primarily written in Python.
How popular is mcpmark?
eval-sys/mcpmark has 433 stars on GitHub.
Where can I find mcpmark?
eval-sys/mcpmark is on GitHub at https://github.com/eval-sys/mcpmark.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.