Language Models

Language Models

underdogs breaking out
01
Neroued/ninfer
+33% /wk +73 ★/dayaccelerating

NInfer is a from-scratch C++/CUDA engine that trades all generality for maximum single-GPU throughput on a closed registry of Qwen checkpoints.

1.6k C++ Inference · Serving · explained
02
xzf-thu/VoiceMem
+28% /wk +51 ★/daysteady

VoiceMem exists because real-time voice agents need to remember who you are, what you said, and how you felt—without making you wait.

1.3k Python Agents · explained
03
syv-ai/qwen38-27b-rtx3090
+26% /wk +47 ★/daysteady

This repo is a surgical stack of vLLM patches, requantization scripts, and speculative decoders built to squeeze Qwen3.8-27B — and up to 268k tokens of context — into a single 24 GB consumer GPU.

1.3k Python Inference · Serving · explained
04
Tencent/WeMM-Embedding
+26% /wk +54 ★/daycooling

It exists to replace a tangle of single-modal encoders with one model that puts text, images, videos, and documents into the same vector space.

1.5k Python RAG · Search · explained
06
fromleda/text-humanizer
+22% /wk +24 ★/daysteady

This tool launders AI-generated text through a chain of translations and LLM rewrites to throw detector tools off the scent.

742 Python LLMOps · Eval · explained
08
KiaBush/persian-text-to-ipa-byt5
+18% /wk +17 ★/daysteady

It exists because Persian TTS and linguistics tools need a modern, learned grapheme-to-phoneme converter that outputs standard IPA without brittle hand-written rules.

651 Python Language Models · explained
09
MakazhanAlpamys/Soup
+15% /wk +128 ★/daycooling

Soup exists because fine-tuning LLMs shouldn't require a cloud budget, SSH, or a PhD in distributed systems.

6k Python ML Frameworks · explained Feature
10
techjarves/Uncensored-Local-Studio
+13% /wk +24 ★/dayaccelerating

It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.

1.3k JavaScript Inference · Serving · explained
11
ximeiorg/Xime
+13% /wk +17 ★/dayaccelerating

Xime is a deliberately minimal, Rime-based Android input method that serves as its author's personal testbed for on-device AI experiments in predictive text and speech recognition.

908 Kotlin Language Models · explained
12
rednote-machine-learning/RedKnot
+12% /wk +42 ★/dayaccelerating

RedKnot accelerates long-context inference by sorting attention heads into four species—global, local, retrieval, and dense—then giving each its own KV reuse strategy and sparsity rules, built as an SGLang extension.

2.4k Python Inference · Serving · explained
13
openJiuwen-ai/jiuwenswarm
+12% /wk +141 ★/daycooling

It breaks complex tasks across a team of specialized LLM agents that refine their own skills as they work.

8.5k Python Agents · explained
14
pguso/rag-from-scratch
+9.5% /wk +22 ★/dayaccelerating

A hands-on Node.js tutorial series that makes you implement embeddings, vector stores, and retrieval yourself so RAG stops feeling like magic.

1.6k JavaScript RAG · Search · explained
15
youssofal/MTPLX
+9.4% /wk +30 ★/daycooling

MTPLX squeezes extra tokens per second out of Apple Silicon by using the multi-token prediction heads that ship with modern models like Qwen 3.6, instead of leaving them idle like most runtimes.

2.2k Python Inference · Serving · explained
16
Sumanth077/Hands-On-AI-Engineering
+8.6% /wk +42 ★/dayaccelerating

Twenty-five bite-sized projects showing how to wire up LLMs, RAG, and agents into things that actually do work.

3.4k Python Learning · explained
17
yukkcat/chatgpt2api
+8.4% /wk +8.3 ★/dayaccelerating

ChatGPT2API exists to reverse-engineer the ChatGPT web interface into a self-hosted OpenAI-compatible API, complete with disposable account pools and a Vue management console.

689 Python Inference · Serving · explained
18
FlashML-org/FreeToken
+7.3% /wk +129 ★/daycooling

Squeezes datacenter-scale MoE inference onto consumer GPUs by treating your desktop’s heterogeneous resources as a unified, elastic runtime.

12.4k Python Inference · Serving · explained
20
e2b-dev/open-computer-use
+6.7% /wk +21 ★/dayaccelerating

It gives open-source LLMs a secure, sandboxed Linux desktop they can click, type, and shell into while you watch.

2.3k Python Agents · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.