LLMOps · Eval

LLMOps · Eval

underdogs · picking up speed
01
AMAP-ML/LongHorizon-Harness
+66% /wk +141 ★/dayaccelerating

It stops long-running agents from drifting by forcing them to verify progress through independent roles before every new round.

1.5k Python Agents · explained Feature
02
wang2122/sprix-sage-router
+65% /wk +336 ★/dayaccelerating

SAGE scores whether an A2A agent should keep a task, recruit complementary teammates, or hand it off entirely—then updates its beliefs from execution evidence.

3.6k Python Agents · explained Feature
04
FailproofAI/failproofai
+49% /wk +187 ★/dayaccelerating

Because coding agents will force-push to main or leak API keys unless something stops them mid-flight.

2.7k MDX Coding Assistants · explained
05
shy3130/tick-stock-panel
+42% /wk +273 ★/dayaccelerating

It exists to give retail A-share traders a lightweight, self-hosted alternative to bloated commercial terminals—no forced cloud data, no pretend AI stock picks.

4.5k Python Domain Apps · explained
06
zhu1090093659/dsh-web
+30% /wk +310 ★/dayaccelerating

A plugin and skin pack that turns the DeepSeek Harness Web UI into a skinnable, pet-equipped control center.

7.3k TypeScript Agents · explained
07
shuxiachai/academic-commercialization-agent
+28% /wk +32 ★/dayaccelerating

It spares tech transfer offices the weeks of manual patent, market, and literature review needed to assess a paper's commercial potential.

806 Python Agents · explained
08
aibox22/readmeX
+27% /wk +22 ★/dayaccelerating

It treats documentation as a batch pipeline: point it at a repo and it emits a README, logo, structure map, and MkDocs wiki—local models optional.

564 Python Coding Assistants · explained
09
typedef-ai/fenic
+22% /wk +21 ★/dayaccelerating

fenic offloads inference-heavy context work into a declarative DataFrame pipeline, then serves the results to any agent framework as bounded, typed tools.

672 Python Agents · explained
10
trailhq/Graft
+25% /wk +250 ★/dayaccelerating

Graft writes a plain-English map of your codebase into linked markdown files so coding agents stop burning tokens rediscovering what they learned last session.

7k TypeScript Coding Assistants · explained
11
java-up-up/nexus-agent
+15% /wk +13 ★/dayaccelerating

It exists because most RAG tutorials end at 'hello vector DB,' while production requires query routing, evidence budgets, and circuit breakers.

592 Java Agents · explained
12
MemTensor/memmy-agent
+32% /wk +83 ★/dayaccelerating

Memmy exists so switching between Cursor, Claude Code, and Codex doesn't mean starting your project history from scratch.

1.8k TypeScript Agents · explained
13
w8123/EnterpriseAgentFramework
+19% /wk +20 ★/dayaccelerating

ReachAI exists because Dify and friends are built for greenfield AI apps, not for wiring agents into existing Java ERP systems that already own the business logic.

721 Java Agents · explained
14
KKKKhazix/human-writing
+18% /wk +66 ★/dayaccelerating

This project exists because AI-generated Chinese is fluently anonymous, so it encodes hard editorial rules and a prose linter to force models to write with the specific gravity of a real person.

2.6k Python Agents · explained
15
Jwuthri/Tracely-ai
+16% /wk +31 ★/dayaccelerating

Tracely converts real production agent failures into hermetic, zero-cost regression tests that block pull requests before they ship the same bug twice.

1.4k Python LLMOps · Eval · explained
16
VectifyAI/OpenKB
+14% /wk +93 ★/dayaccelerating

OpenKB compiles raw documents into a persistent, interlinked wiki so knowledge accumulates instead of being re-derived on every query.

4.5k Python RAG · Search · explained
17
jordan-gibbs/hyperresearch
+16% /wk +48 ★/dayaccelerating

Hyperresearch turns Claude Code into a tiered, multi-agent research pipeline that persists every source in a searchable, compounding vault.

2.2k Python Agents · explained
18
ENTERPILOT/GoModel
+14% /wk +23 ★/dayaccelerating

GoModel exists to spare you from juggling a dozen LLM API formats by unifying them behind a single OpenAI-compatible endpoint written in Go.

1.1k Go Inference · Serving · explained
19
NVIDIA-NeMo/Gym
+14% /wk +23 ★/dayaccelerating

Infrastructure to evaluate and improve agents in reproducible, stateful environments at scale.

1.2k Python Agents · explained
20
deanpeters/product-manager-prompts
+14% /wk +22 ★/dayaccelerating

These prompts exist because treating ChatGPT like a magic template machine is why your user stories still sound like Mad Libs.

1.1k Python LLMOps · Eval · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.