NInfer is a from-scratch C++/CUDA engine that trades all generality for maximum single-GPU throughput on a closed registry of Qwen checkpoints.
Language Models
underdogs · picking up speedBecause serving DeepSeek V4 Flash at 1M-token context across two DGX Spark nodes takes more than a `docker run`.
Xime is a deliberately minimal, Rime-based Android input method that serves as its author's personal testbed for on-device AI experiments in predictive text and speech recognition.
A hands-on Node.js tutorial series that makes you implement embeddings, vector stores, and retrieval yourself so RAG stops feeling like magic.
It gives open-source LLMs a secure, sandboxed Linux desktop they can click, type, and shell into while you watch.
It exists because squeezing 27B-parameter models onto a single consumer GPU requires more than generic kernels and wishful thinking.
It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.
YuE is an open foundation model that turns lyrics into full, multi-minute songs with vocals and accompaniment, offering an open-weight alternative to closed commercial generators.
Moshi is a speech-text foundation model built for real-time, full-duplex conversation, using a custom streaming neural codec that compresses 24 kHz audio to 1.1 kbps.
To deliver a family of tiny language models that squeeze as much capability as possible into edge-friendly checkpoints, with the latest 1B release claiming open-source SOTA in its class and a switchable reasoning mode.
This repo exists so developers can lift working Python patterns for LLMs and agents instead of writing boilerplate from scratch.
Twenty-five bite-sized projects showing how to wire up LLMs, RAG, and agents into things that actually do work.
An AI companion platform that remembers, feels, and stares at your screen—now with a Steam release and a 1000-year SSL certificate.
ChatGPT2API exists to reverse-engineer the ChatGPT web interface into a self-hosted OpenAI-compatible API, complete with disposable account pools and a Vue management console.
RedKnot accelerates long-context inference by sorting attention heads into four species—global, local, retrieval, and dense—then giving each its own KV reuse strategy and sparsity rules, built as an SGLang extension.
It exists to replace 'formula first, API later' with broken experiments that teach you why PPO actually works.
It squeezes massive PyTorch language models into a fraction of their usual memory using 8-bit and 4-bit quantization, enabling inference and fine-tuning on consumer hardware.
A privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.
It exists so you can swap GPT for open-source, speech, and multimodal models by changing a single line of client code.
A minimal VLM you can train from scratch on one GPU in two hours for the price of a coffee.


