Your agent's production traces become its next fine-tune
Overmind closes the loop: trace your agent, score it, fine-tune it, and serve the result — all from one platform you can self-host.

What it does
Overmind is an LLMOps platform that scans your agent’s repo into a graph of capabilities, prompts, and tools, then ingests OpenTelemetry traces from production. Those traces get scored against evals you define, become versioned datasets, and feed fine-tunes that you own — benchmarked against your own production data and served through one OpenAI-compatible API. It’s the full improve-your-agent loop in one box, accessible via a hosted console, CLI, REST API, or an MCP server for Cursor, Claude Code, OpenCode, and Codex.
The interesting bit
The optimiser runs prompt, tool, and control-flow experiments as changes in your repo, and the winning variant lands as a git diff — improvement becomes reviewable code, not a magic number in a dashboard. Also notable: the SDK/CLI are MIT while the platform is AGPL-3.0 and self-hostable, so your traces and training data can stay entirely inside your network.
Key highlights
- Production traces scored on arrival and matched to the capability that produced them
- Datasets built from traces, LLM calls, or uploads — versioned for eval and training
- Fine-tuned models you can download, retrain, or roll back; weights are yours
- MCP integration with per-tool cost declarations (
free,compute,llm,gpu) and no destructive tools - Can ingest existing traces from Langfuse, LangSmith, Braintrust, or Galileo via connectors
Caveats
- Self-hosting is not trivial: the stack needs Postgres, Redis, Celery, plus external dependencies on Modal (training/serving), OpenRouter, AWS S3, and Hugging Face tokens before it will even boot
- The hosted/self-hosted split means the interesting ML work runs on Modal’s infrastructure, not purely “inside your own network”
Verdict
Worth a look if you run agents in production and want a structured path from traces to fine-tunes without building ML infra. If you just want eval dashboards, lighter tools exist — the fine-tuning loop is the whole point here.
Frequently asked
- What is overmind-core/overmind?
- Overmind closes the loop: trace your agent, score it, fine-tune it, and serve the result — all from one platform you can self-host.
- Is overmind open source?
- Yes — overmind-core/overmind is open source, released under the AGPL-3.0 license.
- What language is overmind written in?
- overmind-core/overmind is primarily written in Python.
- How popular is overmind?
- overmind-core/overmind has 521 stars on GitHub.
- Where can I find overmind?
- overmind-core/overmind is on GitHub at https://github.com/overmind-core/overmind.