Inference · Serving

Inference · Serving

big names on the move
01
diegosouzapw/OmniRoute
+1595 ★/dayaccelerating

OmniRoute keeps your coding agents online by automatically failing over across 177 AI providers—including free tiers—when quotas run dry.

30k TypeScript Inference · Serving · explained Feature
02
earendil-works/pi
+733 ★/dayaccelerating

Pi bundles a unified LLM API, agent runtime, and interactive coding CLI into a monorepo that treats every dependency update as a potential attack.

77.6k TypeScript Coding Assistants · explained Feature
04
Wei-Shaw/sub2api
+198 ★/dayaccelerating

Sub2API pools AI subscriptions behind a metered gateway so teams or resellers can distribute API quotas without building their own billing stack.

34.2k Go LLMOps · Eval · explained
05
decolua/9router
+137 ★/dayaccelerating

9Router is a local proxy that auto-switches your AI coding tools from paid to free providers and compresses token-heavy tool outputs so you stop hitting limits.

23.6k JavaScript Coding Assistants · explained
06
cheahjs/free-llm-api-resources
+132 ★/dayaccelerating

It catalogs legitimate services offering free API access to large language models, complete with rate limits, model lists, and data-privacy caveats.

28.2k Python Learning · explained
07
Comfy-Org/ComfyUI
+131 ★/daycooling

It exists because clicking 'generate' isn't enough when you need to control every model, parameter, and preprocessing step.

122.2k Python Image · Video · Audio · explained
08
QuantumNous/new-api
+109 ★/daysteady

Because juggling native APIs from a dozen LLM vendors, each with its own auth and billing, is a recipe for migraines.

43.4k TypeScript Inference · Serving · explained
09
ggml-org/llama.cpp
+104 ★/daycooling

It exists to run large language models on virtually any hardware—from Apple Silicon to RISC-V to your browser—with zero external dependencies and minimal setup.

121.6k C++ Inference · Serving · explained
10
BerriAI/litellm
+103 ★/dayaccelerating

Because swapping from GPT-4o to Claude shouldn't require rewriting your request plumbing.

54.7k Python LLMOps · Eval · explained
11
lyogavin/airllm
+98 ★/daycooling

AirLLM slices giant transformers into layer shards so they fit in consumer VRAM without quantization or distillation.

24k Jupyter Notebook Inference · Serving · explained
12
vllm-project/vllm
+81 ★/daycooling

vLLM is an open-source inference engine that pages attention key-value memory like an operating system to drive higher GPU throughput, then exposes it through an OpenAI-compatible API.

87.2k Python Inference · Serving · explained
13
OpenBMB/VoxCPM
+75 ★/daycooling

VoxCPM2 proves TTS doesn't need discrete tokens: a 2B-parameter diffusion model generates continuous 48kHz speech for 30 languages and text-prompted voice cloning.

34.2k Python Image · Video · Audio · explained
14
unslothai/unsloth
+73 ★/dayaccelerating

It wraps local inference and fine-tuning for open models in a web UI, using custom kernels to squeeze more performance out of desktop GPUs than standard tooling.

68.9k Python Inference · Serving · explained
15
ollama/ollama
+69 ★/daysteady

It exists so you can download, run, and chat with open-weight LLMs locally through one CLI and REST API, keeping inference on your own silicon.

176.9k Go Inference · Serving · explained
16
ggml-org/whisper.cpp
+64 ★/dayaccelerating

A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.

52.3k C++ Image · Video · Audio · explained
17
danny-avila/LibreChat
+52 ★/dayaccelerating

LibreChat bundles every major LLM provider into a single self-hosted chat platform so teams don't have to choose—or leak data.

41.3k TypeScript Chat Assistants · explained
18
sgl-project/sglang
+40 ★/dayaccelerating

SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.

30.7k Python Inference · Serving · explained
19
SillyTavern/SillyTavern
+40 ★/daycooling

SillyTavern gives AI hobbyists a single, local interface to command text generators, image engines, and voice models across dozens of backends without losing control of their prompts or data.

31.1k JavaScript Chat Assistants · explained
20
OpenBB-finance/OpenBB
+38 ★/daycooling

OpenBB normalizes proprietary and public financial data so engineers can feed the same sources to Python scripts, REST APIs, Excel, and AI agents without rebuilding integrations.

71k Python Domain Apps · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.