Inference · Serving

Inference · Serving

underdogs breaking out
01
antirez/h3.c
+189% /wk +433 ★/daysteady

Native C inference for MiniMax-H3 video and audio generation on Apple Silicon, with an interactive terminal UI and an SSD-streaming mode that cuts resident transformer memory from ~36 GiB to ~2 GiB.

1.6k C Inference · Serving · explained
02
duolahypercho/codex-router
+69% /wk +200 ★/dayaccelerating

It exists so you can run DeepSeek and Kimi inside OpenAI’s Codex app without switching editors or leaking API keys into chat.

2k JavaScript Coding Assistants · explained
04
seakee/CPA-Manager-Plus
+45% /wk +163 ★/dayaccelerating

It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.

2.5k Go LLMOps · Eval · explained
05
AtomicBot-ai/atomic-agent
+36% /wk +92 ★/dayaccelerating

It keeps the entire agent loop—prompts, tool calls, browser state, and memory—on your laptop so you don't have to rent a control plane in the cloud.

1.8k TypeScript Agents · explained
06
jegly/Box
+22% /wk +24 ★/dayaccelerating

A privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.

749 Kotlin Inference · Serving · explained
07
espressif/esp-claw
+22% /wk +62 ★/dayaccelerating

Espressif's C framework turns cheap microcontrollers into edge AI agents you program through IM chat.

2k C Agents · explained
08
cactus-compute/needle
+21% /wk +128 ★/dayaccelerating

A distilled Gemini 3.1 that fits on watches and glasses, finetunable on a laptop.

4.3k Python Language Models · explained
09
lidge-jun/opencodex
+20% /wk +277 ★/daycooling

It breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.

9.6k TypeScript Coding Assistants · explained
10
sqliteai/warp
+16% /wk +48 ★/daycooling

WASTE exists to find out how far local inference can be pushed when model weights live mostly on fast storage instead of RAM.

2.1k C Inference · Serving · explained Feature
11
YGYOOO/WorldX
+16% /wk +29 ★/dayaccelerating

WorldX turns one sentence into a self-running simulation of AI agents who gossip, scheme, and remember grudges without a script.

1.3k TypeScript Agents · explained
12
peteonrails/voxtype
+15% /wk +24 ★/dayaccelerating

Because dictating to your cursor on Linux shouldn’t require a browser tab or a cloud subscription.

1.1k Rust Inference · Serving · explained
13
AgnesAI-Labs/AgnesAI-Models
+14% /wk +53 ★/dayaccelerating

Agnes AI is a hosted multimodal API that speaks OpenAI's protocol, letting you reroute existing clients to its text, image, video, and agent models by changing a base URL.

2.6k Inference · Serving · explained
14
astaxie/TokenHub
+14% /wk +20 ★/dayaccelerating

TokenHub exists because dropping a shared OpenAI key into a Slack channel does not scale past one invoice and zero accountability.

997 Go Inference · Serving · explained
16
artokun/comfyui-mcp
+13% /wk +11 ★/dayaccelerating

It turns any LLM into a ComfyUI operator that edits live graphs, manages models, and runs workflows instead of just forwarding prompts.

577 TypeScript Agents · explained
17
superlinked/sie
+13% /wk +52 ★/dayaccelerating

SIE replaces the usual tangle of separate embedding, reranking, and extraction servers with a single open-source container that scales from a laptop to Kubernetes.

2.8k Python RAG · Search · explained
18
hero8152/Infinite-Canvas
+13% /wk +49 ★/dayaccelerating

One desktop UI that wires together ComfyUI, OpenAI, Gemini, ModelScope, and a dozen other generative APIs—plus some very opinionated legal terms.

2.6k Python App Builders · explained
19
drumih/turbo-fieldfare
+13% /wk +109 ★/daycooling

TurboFieldfare streams individual experts from SSD on demand so Gemma 4 26B-A4B fits inside ~2 GB of RAM on any Apple Silicon Mac.

5.9k Swift Inference · Serving · explained
20
foxhui/WebAI2API
+12% /wk +22 ★/dayaccelerating

It automates web-only AI services through a stealth browser to expose them as OpenAI-compatible endpoints with multi-account isolation.

1.2k JavaScript Inference · Serving · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.