SAGE scores whether an A2A agent should keep a task, recruit complementary teammates, or hand it off entirely—then updates its beliefs from execution evidence.
LLMOps · Eval
underdogs · picking up speedIt stops long-running agents from drifting by forcing them to verify progress through independent roles before every new round.
It wraps DeepSeek Harness in a Tauri shell so you can skip the Node.js, pnpm, and Docker chores entirely.
Hyperresearch turns Claude Code into a tiered, multi-agent research pipeline that persists every source in a searchable, compounding vault.
A plugin and skin pack that turns the DeepSeek Harness Web UI into a skinnable, pet-equipped control center.
It spares tech transfer offices the weeks of manual patent, market, and literature review needed to assess a paper's commercial potential.
It exists because most RAG tutorials end at 'hello vector DB,' while production requires query routing, evidence budgets, and circuit breakers.
It treats documentation as a batch pipeline: point it at a repo and it emits a README, logo, structure map, and MkDocs wiki—local models optional.
A local-first desktop workbench that structures the entire novel pipeline—from premise to final draft—so your AI doesn’t lose track of the plot by chapter three.
Graft writes a plain-English map of your codebase into linked markdown files so coding agents stop burning tokens rediscovering what they learned last session.
Memmy exists so switching between Cursor, Claude Code, and Codex doesn't mean starting your project history from scratch.
ClawBench measures whether AI browser agents can handle real-world online tasks—booking flights, ordering food, applying for jobs—on live websites rather than sanitized sandboxes.
AgentInspect gives TypeScript AI agents a local evidence loop to debug runs, enforce CI trajectory rules, and share redacted traces without a cloud account.
It’s a self-hosted LLM gateway that automatically routes to the cheapest capable model among your keys and uses a hosted fallback when you lack coverage.
This project exists because AI-generated Chinese is fluently anonymous, so it encodes hard editorial rules and a prose linter to force models to write with the specific gravity of a real person.
Tracely converts real production agent failures into hermetic, zero-cost regression tests that block pull requests before they ship the same bug twice.
GoModel exists to spare you from juggling a dozen LLM API formats by unifying them behind a single OpenAI-compatible endpoint written in Go.
These prompts exist because treating ChatGPT like a magic template machine is why your user stories still sound like Mad Libs.
Infrastructure to evaluate and improve agents in reproducible, stateful environments at scale.
It exists because running LangGraph agents shouldn’t require a LangSmith enterprise license or a cloud bill.
