The project squeezes a 28.9-million-parameter language model onto an ESP32-S3 microcontroller by keeping most of its weights in slow flash memory and reading only what each token needs.
Language Models
underdogs breaking outA demo repo for running extreme-quantized language models locally without needing a research cluster.
It breaks complex tasks across a team of specialized LLM agents that refine their own skills as they work.
Xime is a deliberately minimal, Rime-based Android input method that serves as its author's personal testbed for on-device AI experiments in predictive text and speech recognition.
TabFM exists so you can run classification and regression on messy, mixed-type tables without retraining a model on your data.
It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.
It split off from `verl` to give diffusion, video, and omni-modality models an RL post-training framework that doesn't treat them like chatbots.
DataFlex stops LLM training loops from wasting compute on static data mixes by dynamically selecting, mixing, and reweighting samples inside LLaMA-Factory.
It turns your local machine into an OpenAI-compatible inference endpoint so agents and IDEs can run on offline models without reconfiguration.
Uses multimodal LLMs to transcribe PDFs into Markdown, preserving complex layouts that traditional extractors mangle.
Curated technical deep-dives covering everything from NVLink signal integrity to Kubernetes GPU scheduling and Huawei NPU porting.
This Go CLI turns a single sentence into a full novel by making Architect, Writer, and Editor LLM agents plan, draft, and review inside a long-loop state machine—no human hand-holding required.
Turns Grok's web interface into a standard API so your existing tools just work.
It exists to run a small army of speech-to-text models through a single GGUF-based ggml runtime that actually checks its math.
Built to prove that hand-written Rust kernels and no framework runtime can serve frontier models without the bloat.
Bytez wraps 175,000+ AI models behind a single endpoint so you don't have to host them yourself.
A curated reading list that maps how the post-training world is moving from static SFT to student-rollout distillation with live teacher feedback.
HRM-Text claims to cut pretraining costs by 130–600× compute and 150–900× data, shipping a full 1B-parameter framework with FSDP2, FlashAttention 3, and a hierarchical recurrent architecture.
OpenMythos is an independent attempt to reconstruct Anthropic’s rumored Claude Mythos architecture as a trainable Recurrent-Depth Transformer with switchable attention and sparse MoE layers.
It’s a macOS dictation app built for people who want real-time Whisper transcription without leaving the keyboard.




