It turns DeepSeek Harness’s command-line Web UI into a native macOS and Windows app so you can skip Node.js setup entirely.
Inference · Serving
big names on the moveOmniRoute keeps your coding agents online by automatically failing over across 177 AI providers—including free tiers—when quotas run dry.
Pi bundles a unified LLM API, agent runtime, and interactive coding CLI into a monorepo that treats every dependency update as a potential attack.
SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.
It aggregates the free tiers of sixteen LLM providers into one OpenAI-compatible endpoint so you can experiment without juggling rate limits.
A Go proxy that handles OAuth login for Claude Code and Codex so you can call them through standard API clients.
9Router is a local proxy that auto-switches your AI coding tools from paid to free providers and compresses token-heavy tool outputs so you stop hitting limits.
It exists because clicking 'generate' isn't enough when you need to control every model, parameter, and preprocessing step.
Sub2API pools AI subscriptions behind a metered gateway so teams or resellers can distribute API quotas without building their own billing stack.
It exists to run large language models on virtually any hardware—from Apple Silicon to RISC-V to your browser—with zero external dependencies and minimal setup.
Because juggling native APIs from a dozen LLM vendors, each with its own auth and billing, is a recipe for migraines.
vLLM is an open-source inference engine that pages attention key-value memory like an operating system to drive higher GPU throughput, then exposes it through an OpenAI-compatible API.
It exists so you can download, run, and chat with open-weight LLMs locally through one CLI and REST API, keeping inference on your own silicon.
A hosted proxy that offers free, rate-limited API access to GPT, DeepSeek, and others for Chinese users who'd rather not tunnel through a VPN.
Because swapping from GPT-4o to Claude shouldn't require rewriting your request plumbing.
AirLLM slices giant transformers into layer shards so they fit in consumer VRAM without quantization or distillation.
It wraps local inference and fine-tuning for open models in a web UI, using custom kernels to squeeze more performance out of desktop GPUs than standard tooling.
It centralizes model definitions so the same architecture works across PyTorch, JAX, vLLM, and llama.cpp without rewrites.
Open-source infrastructure for building, benchmarking, and deploying AI agents that control real macOS, Linux, and Windows desktops.
Odysseus is a self-hosted AI workspace that replicates the ChatGPT experience on your own hardware while also handling email, calendars, documents, and deep research without shipping your data to third-party clouds.

