Magnitude is an open-source agent that bundles its own local-model runner so you can skip the Ollama setup and run fully offline.
Inference · Serving
underdogs · picking up speedIt breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.
NInfer is a from-scratch C++/CUDA engine that trades all generality for maximum single-GPU throughput on a closed registry of Qwen checkpoints.
Real-time screen translation for Android games and manga, no root needed.
A WebGPU library that treats shader files like typed TypeScript modules and runs identically in the browser, headless Node, or a deterministic mock for tests.
Because serving DeepSeek V4 Flash at 1M-token context across two DGX Spark nodes takes more than a `docker run`.
GigaAM exists because Russian call centers, music, and atypical speech deserve a dedicated open-source foundation model instead of hand-me-down multilingual checkpoints.
It gives modern audio models a shared native runtime so you can stop managing Python package conflicts and start generating speech, music, and transcripts locally.
It rounds up pre-built Windows binaries for AI libraries that typically force users into complicated, error-prone source builds.
It provides the theoretically correct causal initialization that autoregressive video distillation was missing, enabling one-step to four-step generation without extra training overhead.
GoModel exists to spare you from juggling a dozen LLM API formats by unifying them behind a single OpenAI-compatible endpoint written in Go.
It packages a C++17 video analytics runtime, a browser-based pipeline editor, and async VLM nodes into a single deployable appliance stack for edge hardware.
It wires Alibaba's open-source Qwen3-TTS into ComfyUI so you can clone, design, and script voices by dragging nodes instead of writing Python.
A deep refactor of one-api that adds subscription billing, real payments, and active-active clustering to a unified LLM gateway.
OpenLake wants storage to bypass the host entirely and land straight in GPU memory.
TypeWhisper is a native macOS app that transcribes speech using local AI models by default, then lets you chain the text through programmable workflows, cloud LLMs, or automation APIs.
It exists because squeezing 27B-parameter models onto a single consumer GPU requires more than generic kernels and wishful thinking.
A training and inference stack that squeezes autoregressive video models onto FP4 weights without making them unwatchable.
A production-hardened fork of slime that keeps massive MoE models from collapsing by obsessing over bit-wise alignment between rollout and training.
It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.



