Inference · Serving

Inference · Serving

underdogs · picking up speed
01
magnitudedev/magnitude
+61% /wk +370 ★/dayaccelerating

Magnitude is an open-source agent that bundles its own local-model runner so you can skip the Ollama setup and run fully offline.

4.3k TypeScript Agents · explained
02
lidge-jun/opencodex
+39% /wk +786 ★/dayaccelerating

It breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.

14.2k TypeScript Coding Assistants · explained
03
Neroued/ninfer
+33% /wk +73 ★/dayaccelerating

NInfer is a from-scratch C++/CUDA engine that trades all generality for maximum single-GPU throughput on a closed registry of Qwen checkpoints.

1.6k C++ Inference · Serving · explained
05
vercel-labs/vgpu
+21% /wk +60 ★/dayaccelerating

A WebGPU library that treats shader files like typed TypeScript modules and runs identically in the browser, headless Node, or a deterministic mock for tests.

2k TypeScript ML Frameworks · explained
07
salute-developers/GigaAM
+20% /wk +22 ★/dayaccelerating

GigaAM exists because Russian call centers, music, and atypical speech deserve a dedicated open-source foundation model instead of hand-me-down multilingual checkpoints.

796 Python Image · Video · Audio · explained
08
0xShug0/audio.cpp
+19% /wk +69 ★/dayaccelerating

It gives modern audio models a shared native runtime so you can stop managing Python package conflicts and start generating speech, music, and transcripts locally.

2.5k C++ Inference · Serving · explained
09
wildminder/AI-windows-whl
+17% /wk +21 ★/dayaccelerating

It rounds up pre-built Windows binaries for AI libraries that typically force users into complicated, error-prone source builds.

859 Inference · Serving · explained
10
thu-ml/Causal-Forcing
+16% /wk +22 ★/dayaccelerating

It provides the theoretically correct causal initialization that autoregressive video distillation was missing, enabling one-step to four-step generation without extra training overhead.

956 Python Image · Video · Audio · explained
11
ENTERPILOT/GoModel
+14% /wk +23 ★/dayaccelerating

GoModel exists to spare you from juggling a dozen LLM API formats by unifying them behind a single OpenAI-compatible endpoint written in Go.

1.1k Go Inference · Serving · explained
12
cosmo-wander-ai/cosmo-edge
+18% /wk +20 ★/dayaccelerating

It packages a C++17 video analytics runtime, a browser-based pipeline editor, and async VLM nodes into a single deployable appliance stack for edge hardware.

797 C Inference · Serving · explained
13
flybirdxx/ComfyUI-Qwen-TTS
+10% /wk +28 ★/dayaccelerating

It wires Alibaba's open-source Qwen3-TTS into ComfyUI so you can clone, design, and script voices by dragging nodes instead of writing Python.

1.9k Python Image · Video · Audio · explained
14
modelbus/one-api-pro
+19% /wk +21 ★/dayaccelerating

A deep refactor of one-api that adds subscription billing, real payments, and active-active clustering to a unified LLM gateway.

770 Go Inference · Serving · explained
16
TypeWhisper/typewhisper-mac
+8.8% /wk +22 ★/dayaccelerating

TypeWhisper is a native macOS app that transcribes speech using local AI models by default, then lets you chain the text through programmable workflows, cloud LLMs, or automation APIs.

1.8k Swift Image · Video · Audio · explained
17
Luce-Org/lucebox
+6.6% /wk +27 ★/dayaccelerating

It exists because squeezing 27B-parameter models onto a single consumer GPU requires more than generic kernels and wishful thinking.

2.8k C++ Inference · Serving · explained
18
NVlabs/LongLive
+6.2% /wk +23 ★/dayaccelerating

A training and inference stack that squeezes autoregressive video models onto FP4 weights without making them unwatchable.

2.6k Python Image · Video · Audio · explained
19
radixark/miles
+12% /wk +48 ★/dayaccelerating

A production-hardened fork of slime that keeps massive MoE models from collapsing by obsessing over bit-wise alignment between rollout and training.

2.8k Python ML Frameworks · explained
20
techjarves/Uncensored-Local-Studio
+13% /wk +24 ★/dayaccelerating

It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.

1.3k JavaScript Inference · Serving · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.