This project exists because AI-generated Chinese is fluently anonymous, so it encodes hard editorial rules and a prose linter to force models to write with the specific gravity of a real person.
LLMOps · Eval
underdogs breaking outRealReplicaBench evaluates whether agents can complete long-horizon business workflows—like listing products or booking freight—instead of just answering questions about them.
QM is an open-source harness that gives every employee an isolated agent workspace while letting teams collaborate in shared channels, without locking the org to a single model.
Raven wraps agents in durable memory and self-refining skills so workflows survive past the chat session.
It exists to detect and optionally block AI agent activity directly on the endpoint, before sensitive actions execute.
Token Monitor reads local logs from two dozen AI coding tools to surface live token burn, costs, and limits in one place, synced across all your machines.
A hybrid CLI tool that uses deterministic pipelines to keep LLM agents from drifting off-target during code review.
Most AI scientist tools are monolithic prompt pipelines; FAROS treats research automation as a composable runtime problem rather than a single-agent stack.
iFixAi runs up to 32 inspections against any LLM or agent and returns a letter-grade scorecard in minutes, using a separate provider as judge so the model isn't grading its own homework.
To wire up public A-share, US, and HK market data into a single local dashboard and let your own AI analyze it without pretending to pick winners.
It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.
It replaces flat vector dumps with a four-tier semantic pyramid and Mermaid symbol graphs so agents remember workflows without drowning in their own tool logs.
Books are too good to leave on the shelf; this system distills them into structured, callable agent skills.
It gives AI dev agents a persistent memory and long-term runtime on your hardware so they learn, schedule tasks, and resume work instead of starting fresh every chat.
A Python layer that makes AI agents explain themselves through structured context graphs, decision trails, and W3C-compliant provenance.
Memmy exists so switching between Cursor, Claude Code, and Codex doesn't mean starting your project history from scratch.
A local-first personal assistant that unpacks the four pillars of agent engineering—harness, loop, memory, and eval—into plain Python you can read in an afternoon.
repowise indexes a codebase into five queryable intelligence layers—dependency graphs, git history, docs, architectural decisions, and deterministic health scores—so MCP-compatible agents can answer "why" instead of grepping for "what".
Coding agents burn tokens re-sending tool schemas, file reads, and history every turn; Paritok sits between your agent and the LLM to compress that bloat non-destructively.
To keep scientific AI local: a desktop workbench where LLMs run Python, R, and query ~80 bio DBs without cloud lock-in.


