Inference · Serving

Inference · Serving

big names · picking up speed
01
lyogavin/airllm
+769 ★/dayaccelerating

AirLLM slices giant transformers into layer shards so they fit in consumer VRAM without quantization or distillation.

30.2k Jupyter Notebook Inference · Serving · explained Feature
02
Comfy-Org/ComfyUI
+254 ★/dayaccelerating

It exists because clicking 'generate' isn't enough when you need to control every model, parameter, and preprocessing step.

124.9k Python Image · Video · Audio · explained
03
sgl-project/sglang
+70 ★/dayaccelerating

SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.

31.5k Python Inference · Serving · explained
05
OpenBB-finance/OpenBB
+51 ★/dayaccelerating

OpenBB normalizes proprietary and public financial data so engineers can feed the same sources to Python scripts, REST APIs, Excel, and AI agents without rebuilding integrations.

71.6k Python Domain Apps · explained
07
ultralytics/ultralytics
+38 ★/dayaccelerating

Ultralytics wants to stop you from stitching together separate repos for every computer vision task by bundling detection, segmentation, tracking, and pose estimation into one YOLO-backed package.

60.4k Python Computer Vision · explained
08
modular/modular
+7.0 ★/dayaccelerating

Modular open-sourced its entire AI stack to let you serve models and write GPU kernels across hardware without switching tools.

26.7k Mojo Inference · Serving · explained
09
qdrant/qdrant
+22 ★/dayaccelerating

Qdrant stores neural network outputs as searchable vectors and lets you filter them with SQL-like payload queries, bridging the gap between embedding models and production search.

33.9k Rust RAG · Search · explained
10
ggml-org/llama.cpp
+113 ★/dayaccelerating

It exists to run large language models on virtually any hardware—from Apple Silicon to RISC-V to your browser—with zero external dependencies and minimal setup.

123.1k C++ Inference · Serving · explained
11
ultralytics/yolov5
+7.6 ★/dayaccelerating

YOLOv5 made real-time object detection as easy as `torch.hub.load`, then exported to everything from iOS to edge chips.

57.8k Python Computer Vision · explained
12
janhq/jan
+16 ★/dayaccelerating

Jan is a desktop chat client that makes running local LLMs as mundane as using ChatGPT, while quietly exposing an OpenAI-compatible API for your other tools.

43.9k TypeScript Chat Assistants · explained
13
google-ai-edge/gallery
+9.1 ★/dayaccelerating

It is a native mobile sandbox for downloading, benchmarking, and interacting with open-source LLMs and multimodal models entirely on-device.

24.4k Kotlin Inference · Serving · explained
14
deepseek-ai/DeepSeek-V3
+10 ★/dayaccelerating

DeepSeek-V3 exists to prove that a 671-billion-parameter model can train end-to-end without a single rollback, activate only 37B parameters per token, and still match leading closed-source systems.

104.1k Python Language Models · explained
15
ray-project/ray
+9.4 ★/dayaccelerating

Ray treats distributed computing as a Python primitive, then layers on libraries for training, tuning, serving, and reinforcement learning.

43.5k Python Inference · Serving · explained
16
karpathy/llm.c
+8.9 ★/dayaccelerating

Because training a transformer shouldn't require 245MB of PyTorch just to multiply matrices.

30.8k Cuda Language Models · explained
17
BerriAI/litellm
+87 ★/dayaccelerating

Because swapping from GPT-4o to Claude shouldn't require rewriting your request plumbing.

55.9k Python LLMOps · Eval · explained
18
danny-avila/LibreChat
+39 ★/dayaccelerating

LibreChat bundles every major LLM provider into a single self-hosted chat platform so teams don't have to choose—or leak data.

41.8k TypeScript Chat Assistants · explained
19
microsoft/BitNet
+2.3 ★/daysteady

Microsoft built an inference engine that lets a single CPU run a 100B-parameter model at human reading speed by using 1.58-bit weights.

39.8k C++ Inference · Serving · explained
20
mozilla-ai/llamafile
+5.1 ★/daysteady

Mozilla wraps llama.cpp and a full model into a single cross-platform executable using an obscure libc trick.

25.5k C++ Inference · Serving · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.