Inference · Serving

Inference · Serving

big names on the move
01
anywhere-labs/dsh-desktop
+1249 ★/dayaccelerating

It turns DeepSeek Harness’s command-line Web UI into a native macOS and Windows app so you can skip Node.js setup entirely.

25.3k TypeScript Agents · explained Feature
02
diegosouzapw/OmniRoute
+481 ★/daycooling

OmniRoute keeps your coding agents online by automatically failing over across 177 AI providers—including free tiers—when quotas run dry.

64.3k TypeScript Inference · Serving · explained Feature
03
earendil-works/pi
+406 ★/dayaccelerating

Pi bundles a unified LLM API, agent runtime, and interactive coding CLI into a monorepo that treats every dependency update as a potential attack.

103.9k TypeScript Coding Assistants · explained Feature
04
sgl-project/sglang
+243 ★/dayaccelerating

SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.

35.8k Python Inference · Serving · explained
05
tashfeenahmed/freellmapi
+210 ★/daycooling

It aggregates the free tiers of sixteen LLM providers into one OpenAI-compatible endpoint so you can experiment without juggling rate limits.

25.4k TypeScript Inference · Serving · explained
07
decolua/9router
+189 ★/dayaccelerating

9Router is a local proxy that auto-switches your AI coding tools from paid to free providers and compresses token-heavy tool outputs so you stop hitting limits.

28.4k JavaScript Coding Assistants · explained
09
Wei-Shaw/sub2api
+117 ★/dayaccelerating

Sub2API pools AI subscriptions behind a metered gateway so teams or resellers can distribute API quotas without building their own billing stack.

41.2k Go LLMOps · Eval · explained
10
ggml-org/llama.cpp
+115 ★/daycooling

It exists to run large language models on virtually any hardware—from Apple Silicon to RISC-V to your browser—with zero external dependencies and minimal setup.

127.8k C++ Inference · Serving · explained
11
QuantumNous/new-api
+89 ★/dayaccelerating

Because juggling native APIs from a dozen LLM vendors, each with its own auth and billing, is a recipe for migraines.

47.8k Go Inference · Serving · explained
12
vllm-project/vllm
+75 ★/daycooling

vLLM is an open-source inference engine that pages attention key-value memory like an operating system to drive higher GPU throughput, then exposes it through an OpenAI-compatible API.

91.4k Python Inference · Serving · explained
13
ollama/ollama
+74 ★/dayaccelerating

It exists so you can download, run, and chat with open-weight LLMs locally through one CLI and REST API, keeping inference on your own silicon.

180.6k Go Inference · Serving · explained
14
chatanywhere/GPT_API_free
+73 ★/dayaccelerating

A hosted proxy that offers free, rate-limited API access to GPT, DeepSeek, and others for Chinese users who'd rather not tunnel through a VPN.

42.3k Inference · Serving · explained
15
BerriAI/litellm
+70 ★/daycooling

Because swapping from GPT-4o to Claude shouldn't require rewriting your request plumbing.

58.5k Python LLMOps · Eval · explained
16
lyogavin/airllm
+64 ★/daycooling

AirLLM slices giant transformers into layer shards so they fit in consumer VRAM without quantization or distillation.

34.1k Jupyter Notebook Inference · Serving · explained Feature
17
unslothai/unsloth
+61 ★/daycooling

It wraps local inference and fine-tuning for open models in a web UI, using custom kernels to squeeze more performance out of desktop GPUs than standard tooling.

76k Python Inference · Serving · explained Feature
18
huggingface/transformers
+47 ★/dayaccelerating

It centralizes model definitions so the same architecture works across PyTorch, JAX, vLLM, and llama.cpp without rewrites.

165.1k Python Language Models · explained
19
trycua/cua
+44 ★/dayaccelerating

Open-source infrastructure for building, benchmarking, and deploying AI agents that control real macOS, Linux, and Windows desktops.

22.5k HTML Agents · explained
20
odysseus-dev/odysseus
+44 ★/daycooling

Odysseus is a self-hosted AI workspace that replicates the ChatGPT experience on your own hardware while also handling email, calendars, documents, and deep research without shipping your data to third-party clouds.

87.1k Python Agents · explained Feature
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.