A ground-up reimplementation of NVIDIA's neural rendering runtime that lets RDNA 3 and 4 cards run DLSS 5 in FSR-enabled DirectX 12 games.
Inference · Serving
underdogs breaking outReef is infrastructure that closes the loop between serving agent traffic, collecting feedback, and shipping updated weights or harnesses without downtime.
Magnitude is an open-source agent that bundles its own local-model runner so you can skip the Ollama setup and run fully offline.
A tutorial that bridges the gap between inference theory and production code by making you build a working mini-sglang first.
It breaks the vendor lock on Codex and Claude Code by translating their API calls to any LLM backend you choose.
NInfer is a from-scratch C++/CUDA engine that trades all generality for maximum single-GPU throughput on a closed registry of Qwen checkpoints.
Real-time screen translation for Android games and manga, no root needed.
This repo is a surgical stack of vLLM patches, requantization scripts, and speculative decoders built to squeeze Qwen3.8-27B — and up to 268k tokens of context — into a single 24 GB consumer GPU.
Because serving DeepSeek V4 Flash at 1M-token context across two DGX Spark nodes takes more than a `docker run`.
A WebGPU library that treats shader files like typed TypeScript modules and runs identically in the browser, headless Node, or a deterministic mock for tests.
GigaAM exists because Russian call centers, music, and atypical speech deserve a dedicated open-source foundation model instead of hand-me-down multilingual checkpoints.
A deep refactor of one-api that adds subscription billing, real payments, and active-active clustering to a unified LLM gateway.
It gives modern audio models a shared native runtime so you can stop managing Python package conflicts and start generating speech, music, and transcripts locally.
It turns a Raspberry Pi 5 into a fully offline kiosk for two people to talk across languages.
It gives generative models a first-person camera and spatial ears, producing synchronized sight and sound as you navigate.
It packages a C++17 video analytics runtime, a browser-based pipeline editor, and async VLM nodes into a single deployable appliance stack for edge hardware.
It rounds up pre-built Windows binaries for AI libraries that typically force users into complicated, error-prone source builds.
Omega-AI was built from scratch in Java so JVM-native developers can train neural nets, run YOLO, and even generate images without bridging into Python ecosystems.
It provides the theoretically correct causal initialization that autoregressive video distillation was missing, enabling one-step to four-step generation without extra training overhead.
A single CLI and credit balance for generating images, video, audio, and text when creator suites are too slow or expensive for automated pipelines.






