Inference · Serving

Inference · Serving

big names · picking up speed
01
diegosouzapw/OmniRoute
+1595 ★/dayaccelerating

OmniRoute keeps your coding agents online by automatically failing over across 177 AI providers—including free tiers—when quotas run dry.

30k TypeScript Inference · Serving · explained Feature
02
earendil-works/pi
+733 ★/dayaccelerating

Pi bundles a unified LLM API, agent runtime, and interactive coding CLI into a monorepo that treats every dependency update as a potential attack.

77.6k TypeScript Coding Assistants · explained Feature
03
cheahjs/free-llm-api-resources
+132 ★/dayaccelerating

It catalogs legitimate services offering free API access to large language models, complete with rate limits, model lists, and data-privacy caveats.

28.2k Python Learning · explained
04
ggml-org/whisper.cpp
+64 ★/dayaccelerating

A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.

52.3k C++ Image · Video · Audio · explained
05
unslothai/unsloth
+73 ★/dayaccelerating

It wraps local inference and fine-tuning for open models in a web UI, using custom kernels to squeeze more performance out of desktop GPUs than standard tooling.

68.9k Python Inference · Serving · explained
06
decolua/9router
+137 ★/dayaccelerating

9Router is a local proxy that auto-switches your AI coding tools from paid to free providers and compresses token-heavy tool outputs so you stop hitting limits.

23.6k JavaScript Coding Assistants · explained
07
chatanywhere/GPT_API_free
+24 ★/dayaccelerating

A hosted proxy that offers free, rate-limited API access to GPT, DeepSeek, and others for Chinese users who'd rather not tunnel through a VPN.

39k Inference · Serving · explained
08
mozilla-ai/llamafile
+12 ★/dayaccelerating

Mozilla wraps llama.cpp and a full model into a single cross-platform executable using an obscure libc trick.

25.4k C++ Inference · Serving · explained
09
huggingface/transformers
+38 ★/dayaccelerating

It centralizes model definitions so the same architecture works across PyTorch, JAX, vLLM, and llama.cpp without rewrites.

163k Python Language Models · explained
10
danny-avila/LibreChat
+52 ★/dayaccelerating

LibreChat bundles every major LLM provider into a single self-hosted chat platform so teams don't have to choose—or leak data.

41.3k TypeScript Chat Assistants · explained
11
BerriAI/litellm
+103 ★/dayaccelerating

Because swapping from GPT-4o to Claude shouldn't require rewriting your request plumbing.

54.7k Python LLMOps · Eval · explained
12
microsoft/BitNet
+8.6 ★/dayaccelerating

Microsoft built an inference engine that lets a single CPU run a 100B-parameter model at human reading speed by using 1.58-bit weights.

39.8k C++ Inference · Serving · explained
13
Wei-Shaw/sub2api
+198 ★/dayaccelerating

Sub2API pools AI subscriptions behind a metered gateway so teams or resellers can distribute API quotas without building their own billing stack.

34.2k Go LLMOps · Eval · explained
15
mudler/LocalAI
+28 ★/dayaccelerating

LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.

47.9k Go Inference · Serving · explained
16
deepseek-ai/DeepSeek-V3
+9.3 ★/dayaccelerating

DeepSeek-V3 exists to prove that a 671-billion-parameter model can train end-to-end without a single rollback, activate only 37B parameters per token, and still match leading closed-source systems.

104k Python Language Models · explained
18
karpathy/llm.c
+9.6 ★/dayaccelerating

Because training a transformer shouldn't require 245MB of PyTorch just to multiply matrices.

30.6k Cuda Language Models · explained
19
google-ai-edge/mediapipe
+17 ★/dayaccelerating

It exists to let developers run customized vision, text, and audio machine learning across mobile, web, and edge hardware without cloud round-trips.

36.3k C++ Computer Vision · explained
20
ultralytics/yolov5
+7.0 ★/dayaccelerating

YOLOv5 made real-time object detection as easy as `torch.hub.load`, then exported to everything from iOS to edge chips.

57.7k Python Computer Vision · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.