It turns DeepSeek Harness’s command-line Web UI into a native macOS and Windows app so you can skip Node.js setup entirely.
Inference · Serving
big names · picking up speed9Router is a local proxy that auto-switches your AI coding tools from paid to free providers and compresses token-heavy tool outputs so you stop hitting limits.
A Go proxy that handles OAuth login for Claude Code and Codex so you can call them through standard API clients.
SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.
Deezer open-sourced its TensorFlow stem splitter so developers can pull vocals, drums, bass, and piano out of a mixed track without training a model from scratch.
Pi bundles a unified LLM API, agent runtime, and interactive coding CLI into a monorepo that treats every dependency update as a potential attack.
It exists so you can compile and deploy large language models to phones, browsers, and nearly any consumer GPU from a single stack.
An open standard that lets you train in PyTorch and deploy on hardware that has never heard of it.
Sub2API pools AI subscriptions behind a metered gateway so teams or resellers can distribute API quotas without building their own billing stack.
Open-source infrastructure for building, benchmarking, and deploying AI agents that control real macOS, Linux, and Windows desktops.
It centralizes model definitions so the same architecture works across PyTorch, JAX, vLLM, and llama.cpp without rewrites.
LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.
A hosted proxy that offers free, rate-limited API access to GPT, DeepSeek, and others for Chinese users who'd rather not tunnel through a VPN.
Because juggling native APIs from a dozen LLM vendors, each with its own auth and billing, is a recipe for migraines.
It unifies two dozen LLM providers behind a single OpenAI-compatible API so you can manage keys, quotas, and load balancing without rewriting client code.
Gradio generates interactive web interfaces from plain Python functions so you can demo models or tools without touching JavaScript.
It exists so you can download, run, and chat with open-weight LLMs locally through one CLI and REST API, keeping inference on your own silicon.
An open-source search engine trying to make relevance the default, not a weekend tuning project.
A reimplementation of OpenAI's Whisper that trades the original inference engine for CTranslate2 and gains up to 4× speed without sacrificing accuracy.
Because training a transformer shouldn't require 245MB of PyTorch just to multiply matrices.



