Language Models

Language Models

underdogs · picking up speed
01
Neroued/ninfer
+33% /wk +73 ★/dayaccelerating

NInfer is a from-scratch C++/CUDA engine that trades all generality for maximum single-GPU throughput on a closed registry of Qwen checkpoints.

1.6k C++ Inference · Serving · explained
03
ximeiorg/Xime
+13% /wk +17 ★/dayaccelerating

Xime is a deliberately minimal, Rime-based Android input method that serves as its author's personal testbed for on-device AI experiments in predictive text and speech recognition.

908 Kotlin Language Models · explained
04
pguso/rag-from-scratch
+9.5% /wk +22 ★/dayaccelerating

A hands-on Node.js tutorial series that makes you implement embeddings, vector stores, and retrieval yourself so RAG stops feeling like magic.

1.6k JavaScript RAG · Search · explained
05
e2b-dev/open-computer-use
+6.7% /wk +21 ★/dayaccelerating

It gives open-source LLMs a secure, sandboxed Linux desktop they can click, type, and shell into while you watch.

2.3k Python Agents · explained
06
Luce-Org/lucebox
+6.6% /wk +27 ★/dayaccelerating

It exists because squeezing 27B-parameter models onto a single consumer GPU requires more than generic kernels and wishful thinking.

2.8k C++ Inference · Serving · explained
07
techjarves/Uncensored-Local-Studio
+13% /wk +24 ★/dayaccelerating

It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.

1.3k JavaScript Inference · Serving · explained
08
multimodal-art-projection/YuE
+5.0% /wk +48 ★/dayaccelerating

YuE is an open foundation model that turns lyrics into full, multi-minute songs with vocals and accompaniment, offering an open-weight alternative to closed commercial generators.

6.6k Python Image · Video · Audio · explained
09
kyutai-labs/moshi
+4.9% /wk +78 ★/dayaccelerating

Moshi is a speech-text foundation model built for real-time, full-duplex conversation, using a custom streaming neural codec that compresses 24 kHz audio to 1.1 kbps.

11k Python Language Models · explained
10
OpenBMB/MiniCPM
+4.8% /wk +74 ★/dayaccelerating

To deliver a family of tiny language models that squeeze as much capability as possible into edge-friendly checkpoints, with the latest 1B release claiming open-source SOTA in its class and a switchable reasoning mode.

10.8k Jupyter Notebook Language Models · explained
11
daveebbelaar/ai-cookbook
+3.8% /wk +24 ★/dayaccelerating

This repo exists so developers can lift working Python patterns for LLMs and agents instead of writing boilerplate from scratch.

4.4k Python Learning · explained
12
Sumanth077/Hands-On-AI-Engineering
+8.6% /wk +42 ★/dayaccelerating

Twenty-five bite-sized projects showing how to wire up LLMs, RAG, and agents into things that actually do work.

3.4k Python Learning · explained
13
Project-N-E-K-O/N.E.K.O
+4.1% /wk +17 ★/dayaccelerating

An AI companion platform that remembers, feels, and stares at your screen—now with a Steam release and a 1000-year SSL certificate.

2.8k Python Agents · explained
14
yukkcat/chatgpt2api
+8.4% /wk +8.3 ★/dayaccelerating

ChatGPT2API exists to reverse-engineer the ChatGPT web interface into a self-hosted OpenAI-compatible API, complete with disposable account pools and a Vue management console.

689 Python Inference · Serving · explained
15
rednote-machine-learning/RedKnot
+12% /wk +42 ★/dayaccelerating

RedKnot accelerates long-context inference by sorting attention heads into four species—global, local, retrieval, and dense—then giving each its own KV reuse strategy and sparsity rules, built as an SGLang extension.

2.4k Python Inference · Serving · explained
16
walkinglabs/hands-on-modern-rl
+3.9% /wk +24 ★/dayaccelerating

It exists to replace 'formula first, API later' with broken experiments that teach you why PPO actually works.

4.3k Python Learning · explained
17
bitsandbytes-foundation/bitsandbytes
+2.0% /wk +24 ★/dayaccelerating

It squeezes massive PyTorch language models into a fraction of their usual memory using 8-bit and 4-bit quantization, enabling inference and fine-tuning on consumer hardware.

8.5k Python Inference · Serving · explained
18
jegly/Box
+3.2% /wk +3.7 ★/dayaccelerating

A privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.

818 Kotlin Inference · Serving · explained
19
xorbitsai/inference
+1.6% /wk +22 ★/dayaccelerating

It exists so you can swap GPT for open-source, speech, and multimodal models by changing a single line of client code.

9.6k Python Inference · Serving · explained
20
jingyaogong/minimind-v
+1.6% /wk +19 ★/dayaccelerating

A minimal VLM you can train from scratch on one GPU in two hours for the price of a coffee.

8.6k Python Language Models · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.