Inference · Serving

Inference · Serving

underdogs breaking out
02
unicity-aos/aos-ce
+53% /wk +557 ★/daysteady

AOS Community Edition is a Rust-based agent operating system that lets agents inspect the runtime, spot missing capabilities, and forge their own least-privilege extensions.

7.3k Rust Agents · explained
03
EvanZhouDev/openai-oauth
+50% /wk +64 ★/dayaccelerating

An unofficial proxy that borrows your ChatGPT/Codex OAuth tokens to serve a local OpenAI-compatible API, bypassing API credit billing.

902 TypeScript Inference · Serving · explained
04
seakee/CPA-Manager-Plus
+36% /wk +112 ★/dayaccelerating

It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.

2.2k TypeScript LLMOps · Eval · explained
05
AtomicBot-ai/atomic-agent
+32% /wk +47 ★/dayaccelerating

It keeps the entire agent loop—prompts, tool calls, browser state, and memory—on your laptop so you don't have to rent a control plane in the cloud.

1k TypeScript Agents · explained
06
routatic/proxy
+21% /wk +27 ★/dayaccelerating

A Go proxy that tricks Claude Code into using $5/month open models through OpenCode instead of Anthropic's API.

892 Go Coding Assistants · explained
07
lidge-jun/opencodex
+19% /wk +129 ★/daysteady

It breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.

4.8k TypeScript Coding Assistants · explained
08
Lynpoint/CyberVerse
+18% /wk +38 ★/dayaccelerating

A self-hosted framework for building real-time voice-first AI agents that persist memory, delegate long tasks to background sub-agents, and optionally show up as lip-synced digital humans.

1.5k Python Agents · explained
09
espressif/esp-claw
+18% /wk +48 ★/dayaccelerating

Espressif's C framework turns cheap microcontrollers into edge AI agents you program through IM chat.

1.9k C Agents · explained
10
astaxie/TokenHub
+15% /wk +14 ★/daysteady

TokenHub exists because dropping a shared OpenAI key into a Slack channel does not scale past one invoice and zero accountability.

650 Go Inference · Serving · explained
11
hero8152/Infinite-Canvas
+14% /wk +48 ★/dayaccelerating

One desktop UI that wires together ComfyUI, OpenAI, Gemini, ModelScope, and a dozen other generative APIs—plus some very opinionated legal terms.

2.4k Python App Builders · explained
12
verl-project/verl-omni
+14% /wk +13 ★/dayaccelerating

It split off from `verl` to give diffusion, video, and omni-modality models an RL post-training framework that doesn't treat them like chatbots.

651 Python ML Frameworks · explained
14
RyanCodrai/turbovec
+12% /wk +242 ★/dayaccelerating

turbovec exists so you can index embeddings immediately—no training, no tuning, no rebuilds—and search them faster than FAISS in a fraction of the RAM.

14.3k Python RAG · Search · explained Feature
15
lidge-jun/ima2-gen
+12% /wk +10 ★/dayaccelerating

It exists because cloud image generators deserve a local memory layer, a branching canvas, and a UI outside the chat thread.

603 TypeScript Image · Video · Audio · explained
16
Osmantic/ODS
+11% /wk +59 ★/daycooling

Dream Server exists because most people would rather pay OpenAI than spend a weekend hand-wiring Docker configs for local LLMs, RAG, and image generation.

3.7k Python Inference · Serving · explained
17
techjarves/Uncensored-Local-Studio
+10% /wk +10 ★/daysteady

It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.

714 JavaScript Inference · Serving · explained
19
Tencent/AngelSlim
+9.2% /wk +20 ★/dayaccelerating

AngelSlim integrates quantization, speculative decoding, and distillation so you can shrink and serve massive models from a single toolkit.

1.5k Python Inference · Serving · explained
20
kvcache-ai/ktransformers
+8.7% /wk +237 ★/dayaccelerating

KTransformers makes CPU-GPU heterogeneous inference and fine-tuning for massive MoE models almost practical on consumer hardware.

19k Python Inference · Serving · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.