Inference · Serving

Inference · Serving

underdogs · picking up speed
01
seakee/CPA-Manager-Plus
+43% /wk +153 ★/dayaccelerating

It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.

2.5k TypeScript LLMOps · Eval · explained
03
jegly/Box
+21% /wk +22 ★/dayaccelerating

A privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.

735 Kotlin Inference · Serving · explained
04
espressif/esp-claw
+20% /wk +56 ★/dayaccelerating

Espressif's C framework turns cheap microcontrollers into edge AI agents you program through IM chat.

1.9k C Agents · explained
05
YGYOOO/WorldX
+15% /wk +27 ★/dayaccelerating

WorldX turns one sentence into a self-running simulation of AI agents who gossip, scheme, and remember grudges without a script.

1.3k TypeScript Agents · explained
06
AgnesAI-Labs/AgnesAI-Models
+14% /wk +53 ★/dayaccelerating

Agnes AI is a hosted multimodal API that speaks OpenAI's protocol, letting you reroute existing clients to its text, image, video, and agent models by changing a base URL.

2.6k Inference · Serving · explained
07
hero8152/Infinite-Canvas
+13% /wk +49 ★/dayaccelerating

One desktop UI that wires together ComfyUI, OpenAI, Gemini, ModelScope, and a dozen other generative APIs—plus some very opinionated legal terms.

2.6k Python App Builders · explained
08
theopenco/llmgateway
+9.9% /wk +21 ★/dayaccelerating

LLM Gateway is an open-source API gateway that normalizes requests to multiple LLM providers behind a single OpenAI-compatible endpoint while tracking token spend and performance.

1.5k TypeScript Inference · Serving · explained
09
AtomicBot-ai/atomic-agent
+32% /wk +76 ★/dayaccelerating

It keeps the entire agent loop—prompts, tool calls, browser state, and memory—on your laptop so you don't have to rent a control plane in the cloud.

1.7k TypeScript Agents · explained
10
superlinked/sie
+11% /wk +43 ★/dayaccelerating

SIE replaces the usual tangle of separate embedding, reranking, and extraction servers with a single open-source container that scales from a laptop to Kubernetes.

2.7k Python RAG · Search · explained
11
ulab-uiuc/LLMRouter
+8.3% /wk +27 ★/dayaccelerating

It matches each prompt to the cheapest capable LLM so you stop paying flagship-model rates for simple queries.

2.3k Python LLMOps · Eval · explained
12
Tavris1/ComfyUI-Easy-Install
+8.2% /wk +20 ★/dayaccelerating

A portable, one-click installer that bundles Python, Git, and dozens of nodes so you can generate images instead of debugging pip.

1.7k Batchfile Inference · Serving · explained
13
nicedreamzapp/claude-code-local
+7.7% /wk +35 ★/dayaccelerating

It tricks Claude Code into talking to local MLX and DeepSeek models instead of Anthropic’s servers, so your code never leaves the Mac.

3.2k Python Coding Assistants · explained
14
basketikun/chatgpt2api
+7.5% /wk +60 ★/dayaccelerating

It turns ChatGPT’s browser-only image generation into a poolable, OpenAI-compatible API so you can self-host programmatic access to GPT-Image-2 and friends.

5.6k Python Inference · Serving · explained
15
Lynpoint/CyberVerse
+6.6% /wk +15 ★/dayaccelerating

A self-hosted framework for building real-time voice-first AI agents that persist memory, delegate long tasks to background sub-agents, and optionally show up as lip-synced digital humans.

1.6k Python Agents · explained
16
gotempsh/temps
+5.7% /wk +5.0 ★/dayaccelerating

Temps exists to compress your deployment platform, error tracker, analytics suite, and AI sandbox provider into one self-hosted Rust binary.

619 Rust LLMOps · Eval · explained
17
lidge-jun/opencodex
+25% /wk +303 ★/dayaccelerating

It breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.

8.6k TypeScript Coding Assistants · explained
18
routatic/proxy
+4.7% /wk +6.1 ★/dayaccelerating

A Go proxy that tricks Claude Code into using $5/month open models through OpenCode instead of Anthropic's API.

920 Go Coding Assistants · explained
19
thu-ml/SageAttention
+4.2% /wk +21 ★/dayaccelerating

A family of drop-in CUDA kernels that quantize transformer attention to INT8 and FP8 to accelerate inference on modern NVIDIA GPUs while claiming no end-to-end quality loss.

3.6k Cuda Inference · Serving · explained
20
Sophomoresty/gemini-web2api
+14% /wk +51 ★/dayaccelerating

It reverse-engineers Google’s web StreamGenerate protocol so any OpenAI client can chat with Gemini for free, without an API key.

2.6k JavaScript Inference · Serving · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.