Inference · Serving

Inference · Serving

big names · picking up speed
01
anywhere-labs/dsh-desktop
+1249 ★/dayaccelerating

It turns DeepSeek Harness’s command-line Web UI into a native macOS and Windows app so you can skip Node.js setup entirely.

25.3k TypeScript Agents · explained Feature
02
decolua/9router
+189 ★/dayaccelerating

9Router is a local proxy that auto-switches your AI coding tools from paid to free providers and compresses token-heavy tool outputs so you stop hitting limits.

28.4k JavaScript Coding Assistants · explained
04
sgl-project/sglang
+243 ★/dayaccelerating

SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.

35.8k Python Inference · Serving · explained
05
deezer/spleeter
+22 ★/dayaccelerating

Deezer open-sourced its TensorFlow stem splitter so developers can pull vocals, drums, bass, and piano out of a mixed track without training a model from scratch.

28.4k Python Image · Video · Audio · explained
06
earendil-works/pi
+406 ★/dayaccelerating

Pi bundles a unified LLM API, agent runtime, and interactive coding CLI into a monorepo that treats every dependency update as a potential attack.

103.9k TypeScript Coding Assistants · explained Feature
07
mlc-ai/mlc-llm
+22 ★/dayaccelerating

It exists so you can compile and deploy large language models to phones, browsers, and nearly any consumer GPU from a single stack.

23.1k Python Inference · Serving · explained
08
onnx/onnx
+19 ★/dayaccelerating

An open standard that lets you train in PyTorch and deploy on hardware that has never heard of it.

21.5k Python Inference · Serving · explained
09
Wei-Shaw/sub2api
+117 ★/dayaccelerating

Sub2API pools AI subscriptions behind a metered gateway so teams or resellers can distribute API quotas without building their own billing stack.

41.2k Go LLMOps · Eval · explained
10
trycua/cua
+44 ★/dayaccelerating

Open-source infrastructure for building, benchmarking, and deploying AI agents that control real macOS, Linux, and Windows desktops.

22.5k HTML Agents · explained
11
huggingface/transformers
+47 ★/dayaccelerating

It centralizes model definitions so the same architecture works across PyTorch, JAX, vLLM, and llama.cpp without rewrites.

165.1k Python Language Models · explained
12
mudler/LocalAI
+28 ★/dayaccelerating

LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.

49k Go Inference · Serving · explained
13
chatanywhere/GPT_API_free
+73 ★/dayaccelerating

A hosted proxy that offers free, rate-limited API access to GPT, DeepSeek, and others for Chinese users who'd rather not tunnel through a VPN.

42.3k Inference · Serving · explained
14
QuantumNous/new-api
+89 ★/dayaccelerating

Because juggling native APIs from a dozen LLM vendors, each with its own auth and billing, is a recipe for migraines.

47.8k Go Inference · Serving · explained
15
songquanpeng/one-api
+18 ★/dayaccelerating

It unifies two dozen LLM providers behind a single OpenAI-compatible API so you can manage keys, quotas, and load balancing without rewriting client code.

36.8k JavaScript Inference · Serving · explained
16
gradio-app/gradio
+7.0 ★/dayaccelerating

Gradio generates interactive web interfaces from plain Python functions so you can demo models or tools without touching JavaScript.

43.5k Python App Builders · explained
17
ollama/ollama
+74 ★/dayaccelerating

It exists so you can download, run, and chat with open-weight LLMs locally through one CLI and REST API, keeping inference on your own silicon.

180.6k Go Inference · Serving · explained
18
meilisearch/meilisearch
+11 ★/dayaccelerating

An open-source search engine trying to make relevance the default, not a weekend tuning project.

59.3k Rust RAG · Search · explained
19
SYSTRAN/faster-whisper
+16 ★/dayaccelerating

A reimplementation of OpenAI's Whisper that trades the original inference engine for CTranslate2 and gains up to 4× speed without sacrificing accuracy.

25.3k Python Inference · Serving · explained
20
karpathy/llm.c
+6.4 ★/dayaccelerating

Because training a transformer shouldn't require 245MB of PyTorch just to multiply matrices.

31k Cuda Language Models · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.