It exists to stop your AI gateway from quietly burning through quotas, cash, and expired OAuth tokens without leaving a paper trail.
Inference · Serving
underdogs · picking up speedIt publishes full open-hardware blueprints for building multi-GPU AI workstations that keep your weights local and your prompts private.
A privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.
Espressif's C framework turns cheap microcontrollers into edge AI agents you program through IM chat.
WorldX turns one sentence into a self-running simulation of AI agents who gossip, scheme, and remember grudges without a script.
Agnes AI is a hosted multimodal API that speaks OpenAI's protocol, letting you reroute existing clients to its text, image, video, and agent models by changing a base URL.
One desktop UI that wires together ComfyUI, OpenAI, Gemini, ModelScope, and a dozen other generative APIs—plus some very opinionated legal terms.
LLM Gateway is an open-source API gateway that normalizes requests to multiple LLM providers behind a single OpenAI-compatible endpoint while tracking token spend and performance.
It keeps the entire agent loop—prompts, tool calls, browser state, and memory—on your laptop so you don't have to rent a control plane in the cloud.
SIE replaces the usual tangle of separate embedding, reranking, and extraction servers with a single open-source container that scales from a laptop to Kubernetes.
It matches each prompt to the cheapest capable LLM so you stop paying flagship-model rates for simple queries.
A portable, one-click installer that bundles Python, Git, and dozens of nodes so you can generate images instead of debugging pip.
It tricks Claude Code into talking to local MLX and DeepSeek models instead of Anthropic’s servers, so your code never leaves the Mac.
It turns ChatGPT’s browser-only image generation into a poolable, OpenAI-compatible API so you can self-host programmatic access to GPT-Image-2 and friends.
A self-hosted framework for building real-time voice-first AI agents that persist memory, delegate long tasks to background sub-agents, and optionally show up as lip-synced digital humans.
Temps exists to compress your deployment platform, error tracker, analytics suite, and AI sandbox provider into one self-hosted Rust binary.
It breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.
A Go proxy that tricks Claude Code into using $5/month open models through OpenCode instead of Anthropic's API.
A family of drop-in CUDA kernels that quantize transformer attention to INT8 and FP8 to accelerate inference on modern NVIDIA GPUs while claiming no end-to-end quality loss.
It reverse-engineers Google’s web StreamGenerate protocol so any OpenAI client can chat with Gemini for free, without an API key.



