Inference · Serving

Inference · Serving

underdogs breaking out
02
Human-Agent-Society/reef
+64% /wk +84 ★/daysteady

Reef is infrastructure that closes the loop between serving agent traffic, collecting feedback, and shipping updated weights or harnesses without downtime.

922 Python Agents · explained
03
magnitudedev/magnitude
+61% /wk +370 ★/dayaccelerating

Magnitude is an open-source agent that bundles its own local-model runner so you can skip the Ollama setup and run fully offline.

4.3k TypeScript Agents · explained
04
datawhalechina/zero-to-sglang
+47% /wk +46 ★/daysteady

A tutorial that bridges the gap between inference theory and production code by making you build a working mini-sglang first.

682 Python Learning · explained
05
lidge-jun/opencodex
+39% /wk +786 ★/dayaccelerating

It breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.

14.2k TypeScript Coding Assistants · explained
06
Neroued/ninfer
+33% /wk +73 ★/dayaccelerating

NInfer is a from-scratch C++/CUDA engine that trades all generality for maximum single-GPU throughput on a closed registry of Qwen checkpoints.

1.6k C++ Inference · Serving · explained
08
syv-ai/qwen38-27b-rtx3090
+26% /wk +47 ★/daysteady

This repo is a surgical stack of vLLM patches, requantization scripts, and speculative decoders built to squeeze Qwen3.8-27B — and up to 268k tokens of context — into a single 24 GB consumer GPU.

1.3k Python Inference · Serving · explained
10
vercel-labs/vgpu
+21% /wk +60 ★/dayaccelerating

A WebGPU library that treats shader files like typed TypeScript modules and runs identically in the browser, headless Node, or a deterministic mock for tests.

2k TypeScript ML Frameworks · explained
11
salute-developers/GigaAM
+20% /wk +22 ★/dayaccelerating

GigaAM exists because Russian call centers, music, and atypical speech deserve a dedicated open-source foundation model instead of hand-me-down multilingual checkpoints.

796 Python Image · Video · Audio · explained
12
modelbus/one-api-pro
+19% /wk +21 ★/dayaccelerating

A deep refactor of one-api that adds subscription billing, real payments, and active-active clustering to a unified LLM gateway.

770 Go Inference · Serving · explained
13
0xShug0/audio.cpp
+19% /wk +69 ★/dayaccelerating

It gives modern audio models a shared native runtime so you can stop managing Python package conflicts and start generating speech, music, and transcripts locally.

2.5k C++ Inference · Serving · explained
15
NoizAI/HelixWorld
+18% /wk +15 ★/daysteady

It gives generative models a first-person camera and spatial ears, producing synchronized sight and sound as you navigate.

601 Python Image · Video · Audio · explained
16
cosmo-wander-ai/cosmo-edge
+18% /wk +20 ★/dayaccelerating

It packages a C++17 video analytics runtime, a browser-based pipeline editor, and async VLM nodes into a single deployable appliance stack for edge hardware.

797 C Inference · Serving · explained
17
wildminder/AI-windows-whl
+17% /wk +21 ★/dayaccelerating

It rounds up pre-built Windows binaries for AI libraries that typically force users into complicated, error-prone source builds.

859 Inference · Serving · explained
18
dromara/Omega-AI
+17% /wk +20 ★/daycooling

Omega-AI was built from scratch in Java so JVM-native developers can train neural nets, run YOLO, and even generate images without bridging into Python ecosystems.

794 Java ML Frameworks · explained
19
thu-ml/Causal-Forcing
+16% /wk +22 ★/dayaccelerating

It provides the theoretically correct causal initialization that autoregressive video distillation was missing, enabling one-step to four-step generation without extra training overhead.

956 Python Image · Video · Audio · explained
20
flatkey-ai/flatkey-cli
+14% /wk +18 ★/daycooling

A single CLI and credit balance for generating images, video, audio, and text when creator suites are too slow or expensive for automated pipelines.

899 JavaScript Creative · Design · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.