LLMOps · Eval

LLMOps · Eval

underdogs · picking up speed
01
wang2122/sprix-sage-router
+66% /wk +349 ★/dayaccelerating

SAGE scores whether an A2A agent should keep a task, recruit complementary teammates, or hand it off entirely—then updates its beliefs from execution evidence.

3.7k Python Agents · explained Feature
02
AMAP-ML/LongHorizon-Harness
+66% /wk +142 ★/dayaccelerating

It stops long-running agents from drifting by forcing them to verify progress through independent roles before every new round.

1.5k Python Agents · explained Feature
04
jordan-gibbs/hyperresearch
+40% /wk +174 ★/dayaccelerating

Hyperresearch turns Claude Code into a tiered, multi-agent research pipeline that persists every source in a searchable, compounding vault.

3k Python Agents · explained
05
zhu1090093659/dsh-web
+31% /wk +329 ★/dayaccelerating

A plugin and skin pack that turns the DeepSeek Harness Web UI into a skinnable, pet-equipped control center.

7.5k TypeScript Agents · explained
06
shuxiachai/academic-commercialization-agent
+28% /wk +32 ★/dayaccelerating

It spares tech transfer offices the weeks of manual patent, market, and literature review needed to assess a paper's commercial potential.

808 Python Agents · explained
07
java-up-up/nexus-agent
+19% /wk +16 ★/dayaccelerating

It exists because most RAG tutorials end at 'hello vector DB,' while production requires query routing, evidence budgets, and circuit breakers.

619 Java Agents · explained
08
aibox22/readmeX
+27% /wk +22 ★/dayaccelerating

It treats documentation as a batch pipeline: point it at a repo and it emits a README, logo, structure map, and MkDocs wiki—local models optional.

568 Python Coding Assistants · explained
09
EthanYoQ/AI-Novel-Writer
+21% /wk +25 ★/dayaccelerating

A local-first desktop workbench that structures the entire novel pipeline—from premise to final draft—so your AI doesn’t lose track of the plot by chapter three.

815 TypeScript App Builders · explained
10
trailhq/Graft
+28% /wk +295 ★/dayaccelerating

Graft writes a plain-English map of your codebase into linked markdown files so coding agents stop burning tokens rediscovering what they learned last session.

7.3k TypeScript Coding Assistants · explained
11
MemTensor/memmy-agent
+35% /wk +93 ★/dayaccelerating

Memmy exists so switching between Cursor, Claude Code, and Codex doesn't mean starting your project history from scratch.

1.9k TypeScript Agents · explained
12
TIGER-AI-Lab/ClawBench
+22% /wk +22 ★/dayaccelerating

ClawBench measures whether AI browser agents can handle real-world online tasks—booking flights, ordering food, applying for jobs—on live websites rather than sanitized sandboxes.

725 Python LLMOps · Eval · explained
13
rajudandigam/agent-inspect
+15% /wk +13 ★/dayaccelerating

AgentInspect gives TypeScript AI agents a local evidence loop to debug runs, enforce CI trajectory rules, and share redacted traces without a cloud account.

594 TypeScript LLMOps · Eval · explained
14
Continuum-AI-Corp/OrcaRouter-Lite
+23% /wk +27 ★/dayaccelerating

It’s a self-hosted LLM gateway that automatically routes to the cheapest capable model among your keys and uses a hosted fallback when you lack coverage.

844 Python Inference · Serving · explained
15
KKKKhazix/human-writing
+18% /wk +66 ★/dayaccelerating

This project exists because AI-generated Chinese is fluently anonymous, so it encodes hard editorial rules and a prose linter to force models to write with the specific gravity of a real person.

2.6k Python Agents · explained
16
Jwuthri/Tracely-ai
+15% /wk +31 ★/dayaccelerating

Tracely converts real production agent failures into hermetic, zero-cost regression tests that block pull requests before they ship the same bug twice.

1.4k Python LLMOps · Eval · explained
17
ENTERPILOT/GoModel
+15% /wk +24 ★/dayaccelerating

GoModel exists to spare you from juggling a dozen LLM API formats by unifying them behind a single OpenAI-compatible endpoint written in Go.

1.2k Go Inference · Serving · explained
18
deanpeters/product-manager-prompts
+14% /wk +23 ★/dayaccelerating

These prompts exist because treating ChatGPT like a magic template machine is why your user stories still sound like Mad Libs.

1.1k Python LLMOps · Eval · explained
19
NVIDIA-NeMo/Gym
+14% /wk +24 ★/dayaccelerating

Infrastructure to evaluate and improve agents in reproducible, stateful environments at scale.

1.2k Python Agents · explained
20
aegra/aegra
+14% /wk +24 ★/dayaccelerating

It exists because running LangGraph agents shouldn’t require a LangSmith enterprise license or a cloud bill.

1.2k Python Agents · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.