Raven wraps agents in durable memory and self-refining skills so workflows survive past the chat session.
LLMOps · Eval
underdogs · picking up speedToken Monitor reads local logs from two dozen AI coding tools to surface live token burn, costs, and limits in one place, synced across all your machines.
A hybrid CLI tool that uses deterministic pipelines to keep LLM agents from drifting off-target during code review.
Most AI scientist tools are monolithic prompt pipelines; FAROS treats research automation as a composable runtime problem rather than a single-agent stack.
It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.
To wire up public A-share, US, and HK market data into a single local dashboard and let your own AI analyze it without pretending to pick winners.
It replaces flat vector dumps with a four-tier semantic pyramid and Mermaid symbol graphs so agents remember workflows without drowning in their own tool logs.
Books are too good to leave on the shelf; this system distills them into structured, callable agent skills.
A Python layer that makes AI agents explain themselves through structured context graphs, decision trails, and W3C-compliant provenance.
repowise indexes a codebase into five queryable intelligence layers—dependency graphs, git history, docs, architectural decisions, and deterministic health scores—so MCP-compatible agents can answer "why" instead of grepping for "what".
It chains a dozen LLM agents into a visual workflow so your on-call engineer only has to tap 'approve' instead of SSHing in at 3 AM.
An open-source audit engine that scores how likely ChatGPT, Perplexity, and Gemini are to cite your site—and generates the fixes.
A cross-CLI skill that turns your static notes into a self-updating knowledge base for Claude, Codex, Gemini, and OpenCode.
DeepSWE measures whether frontier coding agents can complete real, long-horizon engineering tasks from active open-source repositories—not just generate snippets, but ship verifiable patches.
Chinese AI output too often reads like a blend of corporate press release and therapy session; this skill file teaches models to cut the performance and keep the facts.
It packages a professor's decade of SIGMOD and NeurIPS experience into structured AI skills, bridging the 'last-mile' gap where generic guides and busy advisors leave grad students stranded.
To keep scientific AI local: a desktop workbench where LLMs run Python, R, and query ~80 bio DBs without cloud lock-in.
Reusable instruction modules that give AI agents structured playbooks for repeatable coding, research, and ops tasks.
iFixAi runs up to 32 inspections against any LLM or agent and returns a letter-grade scorecard in minutes, using a separate provider as judge so the model isn't grading its own homework.
VulnClaw exists so that a single natural-language sentence can trigger the entire reconnaissance-to-report pipeline without manually orchestrating a dozen separate tools.




