A skill that spawns parallel reasoning processes under distorted cognitive frames, then scores and prunes them with a separate critic pass.
LLMOps · Eval
underdogs breaking outAn agent skill that forces you to approve the visual metaphor and static layout before it spends your Gemini credits on video generation.
Raven wraps agents in durable memory and self-refining skills so workflows survive past the chat session.
iFixAi runs up to 32 inspections against any LLM or agent and returns a letter-grade scorecard in minutes, using a separate provider as judge so the model isn't grading its own homework.
codex-keysmith exists because copying a Markdown file and editing one TOML key is simple, but existing files, active hooks, and interrupted writes are not.
A local-first personal assistant that unpacks the four pillars of agent engineering—harness, loop, memory, and eval—into plain Python you can read in an afternoon.
Hyperresearch turns Claude Code into a tiered, multi-agent research pipeline that persists every source in a searchable, compounding vault.
pilotfish keeps Claude Fable 5 in the boardroom by delegating the grunt work to cheaper Sonnet and Haiku subagents.
It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.
Token Monitor reads local logs from two dozen AI coding tools to surface live token burn, costs, and limits in one place, synced across all your machines.
Moss exists because calling out to a remote vector database adds 200–500 ms of latency—enough to kill a real-time conversation—so it runs embedding and search inside your process instead.
PhyAgentOS treats robot hardware like pluggable drivers so the same agentic session can run in simulation or on a real arm without rewriting the control stack.
AIHelms wraps LiteLLM in a Vue management layer so finance can trace every token back to the department that spent it.
A hybrid CLI tool that uses deterministic pipelines to keep LLM agents from drifting off-target during code review.
A single Zig binary with an embedded Svelte dashboard that installs, supervises, and cross-wires local AI agents, workflow engines, and tracing tools so you don't have to juggle separate terminals.
Books are too good to leave on the shelf; this system distills them into structured, callable agent skills.
repowise indexes a codebase into five queryable intelligence layers—dependency graphs, git history, docs, architectural decisions, and deterministic health scores—so MCP-compatible agents can answer "why" instead of grepping for "what".
Pairs a 424-page textbook with Jupyter notebooks to teach AI agent design patterns.
It chains a dozen LLM agents into a visual workflow so your on-call engineer only has to tap 'approve' instead of SSHing in at 3 AM.
Agentlas OS compiles plain-language requests into portable, ownable agent packages that run locally across whatever LLM host you already use, while a Hub and owner-scoped Cloud handle sharing and retrieval.




