hamelsmu/evals-skills
A plugin of AI evaluation skills that guides AI coding agents to audit, diagnose, and improve LLM evaluation pipelines.

This repository provides a collection of skills designed to guide AI coding agents in building and auditing LLM evaluation pipelines. It includes skills like eval-audit to surface common problems in evaluations and error-analysis to help categorize failures from traces. The skills are distributed as a plugin for Claude Code and as a standalone CLI tool, allowing developers to integrate evaluation guidance directly into their AI assistant workflows.
Frequently asked
- What is hamelsmu/evals-skills?
- A plugin of AI evaluation skills that guides AI coding agents to audit, diagnose, and improve LLM evaluation pipelines.
- Is evals-skills open source?
- Yes — hamelsmu/evals-skills is open source, released under the MIT license.
- How popular is evals-skills?
- hamelsmu/evals-skills has 1.4k stars on GitHub.
- Where can I find evals-skills?
- hamelsmu/evals-skills is on GitHub at https://github.com/hamelsmu/evals-skills.