LLMOps · Eval

LLMOps · Eval

underdogs breaking out
01
KKKKhazix/human-writing
+346% /wk +790 ★/daysteady

This project exists because AI-generated Chinese is fluently anonymous, so it encodes hard editorial rules and a prose linter to force models to write with the specific gravity of a real person.

1.6k Python Agents · explained
02
Accio-org/RealReplicaBench
+144% /wk +214 ★/daysteady

RealReplicaBench evaluates whether agents can complete long-horizon business workflows—like listing products or booking freight—instead of just answering questions about them.

1k HTML LLMOps · Eval · explained
03
yc-software/qm
+85% /wk +1463 ★/daysteady

QM is an open-source harness that gives every employee an isolated agent workspace while letting teams collaborate in shared channels, without locking the org to a single model.

12.1k TypeScript Agents · explained
04
EverMind-AI/Raven
+61% /wk +305 ★/dayaccelerating

Raven wraps agents in durable memory and self-refining skills so workflows survive past the chat session.

3.5k Python Agents · explained
05
perplexityai/numbat
+51% /wk +53 ★/daysteady

It exists to detect and optionally block AI agent activity directly on the endpoint, before sensitive actions execute.

734 Go LLMOps · Eval · explained
06
Javis603/token-monitor
+47% /wk +80 ★/dayaccelerating

Token Monitor reads local logs from two dozen AI coding tools to surface live token burn, costs, and limits in one place, synced across all your machines.

1.2k JavaScript LLMOps · Eval · explained
08
OpenNSWM-Lab/FAROS
+47% /wk +209 ★/dayaccelerating

Most AI scientist tools are monolithic prompt pipelines; FAROS treats research automation as a composable runtime problem rather than a single-agent stack.

3.1k Python Agents · explained
09
ifixai-ai/iFixAi
+46% /wk +415 ★/dayaccelerating

iFixAi runs up to 32 inspections against any LLM or agent and returns a letter-grade scorecard in minutes, using a separate provider as judge so the model isn't grading its own homework.

6.3k Python LLMOps · Eval · explained
10
simonlin1212/Vibe-Research
+43% /wk +116 ★/dayaccelerating

To wire up public A-share, US, and HK market data into a single local dashboard and let your own AI analyze it without pretending to pick winners.

1.9k TypeScript Domain Apps · explained
11
seakee/CPA-Manager-Plus
+43% /wk +148 ★/dayaccelerating

It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.

2.4k TypeScript LLMOps · Eval · explained
12
TencentCloud/TencentDB-Agent-Memory
+41% /wk +968 ★/dayaccelerating

It replaces flat vector dumps with a four-tier semantic pyramid and Mermaid symbol graphs so agents remember workflows without drowning in their own tool logs.

16.4k TypeScript Agents · explained
13
kangarooking/cangjie-skill
+37% /wk +347 ★/dayaccelerating

Books are too good to leave on the shelf; this system distills them into structured, callable agent skills.

6.5k Python Agents · explained Feature
14
kirodotdev/KiroCrew
+33% /wk +89 ★/daysteady

It gives AI dev agents a persistent memory and long-term runtime on your hardware so they learn, schedule tasks, and resume work instead of starting fresh every chat.

1.9k Python Agents · explained
15
semantica-agi/semantica
+33% /wk +99 ★/dayaccelerating

A Python layer that makes AI agents explain themselves through structured context graphs, decision trails, and W3C-compliant provenance.

2.1k Python Agents · explained
16
MemTensor/memmy-agent
+32% /wk +27 ★/daysteady

Memmy exists so switching between Cursor, Claude Code, and Codex doesn't mean starting your project history from scratch.

593 TypeScript Agents · explained
17
ShenSeanChen/waku-agent
+32% /wk +42 ★/dayaccelerating

A local-first personal assistant that unpacks the four pillars of agent engineering—harness, loop, memory, and eval—into plain Python you can read in an afternoon.

931 Python Agents · explained
18
repowise-dev/repowise
+31% /wk +215 ★/dayaccelerating

repowise indexes a codebase into five queryable intelligence layers—dependency graphs, git history, docs, architectural decisions, and deterministic health scores—so MCP-compatible agents can answer "why" instead of grepping for "what".

4.8k Python Coding Assistants · explained
19
Paritok-official/paritok-4b-v1
+31% /wk +25 ★/daysteady

Coding agents burn tokens re-sending tool schemas, file reads, and history every turn; Paritok sits between your agent and the LLM to compress that bloat non-destructively.

577 Python LLMOps · Eval · explained
20
xuzhougeng/wisp-science
+29% /wk +40 ★/dayaccelerating

To keep scientific AI local: a desktop workbench where LLMs run Python, R, and query ~80 bio DBs without cloud lock-in.

981 HTML Agents · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.