Image · Video · Audio

Image · Video · Audio

big names · picking up speed
01
debpalash/VoiceStudio
+1362 ★/dayaccelerating

OmniVoice Studio bundles voice cloning, dubbing, dictation, and TTS into a desktop app that keeps all audio processing off the internet and away from API keys.

22.3k Python Image · Video · Audio · explained Feature
02
sgl-project/sglang
+243 ★/dayaccelerating

SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.

35.8k Python Inference · Serving · explained
03
Anil-matcha/Open-Generative-AI
+83 ★/dayaccelerating

An Electron app that wraps 200+ generative models behind a single UI, with an unusual pitch: no guardrails, no cloud lock-in, and a split personality between local and remote inference.

28.3k JavaScript Image · Video · Audio · explained
04
deezer/spleeter
+22 ★/dayaccelerating

Deezer open-sourced its TensorFlow stem splitter so developers can pull vocals, drums, bass, and piano out of a mixed track without training a model from scratch.

28.4k Python Image · Video · Audio · explained
05
openai/whisper
+70 ★/dayaccelerating

To give developers a single, general-purpose speech model that handles transcription, translation, and language identification by treating tasks as tokens to predict.

108.9k Python Image · Video · Audio · explained
06
jamiepine/voicebox
+99 ★/dayaccelerating

Voicebox is a local-first alternative to ElevenLabs and WisprFlow that clones voices, dictates anywhere, and gives AI agents a mouthpiece — all without shipping audio to the cloud.

52.9k TypeScript Image · Video · Audio · explained Feature
07
microsoft/VibeVoice
+60 ★/dayaccelerating

VibeVoice is a family of open-source speech models from Microsoft built to ingest, transcribe, and generate very long audio sessions—up to an hour—in a single pass.

54k Python Image · Video · Audio · explained
08
mudler/LocalAI
+28 ★/dayaccelerating

LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.

49k Go Inference · Serving · explained
10
fishaudio/fish-speech
+16 ★/dayaccelerating

Fish Speech S2 Pro is a 4B-parameter dual-autoregressive TTS model that treats inline emotional tags like `[whisper]` as native tokens to synthesize speech in over 80 languages.

32.6k Python Image · Video · Audio · explained
11
SYSTRAN/faster-whisper
+16 ★/dayaccelerating

A reimplementation of OpenAI's Whisper that trades the original inference engine for CTranslate2 and gains up to 4× speed without sacrificing accuracy.

25.3k Python Inference · Serving · explained
13
OpenBMB/MiniCPM-V
+7.9 ★/dayaccelerating

To squeeze multimodal understanding—vision, video, and even real-time speech—into models small enough to run natively on a handset.

26.3k Python Image · Video · Audio · explained
14
modelscope/FunASR
+16 ★/dayaccelerating

A Chinese speech toolkit that bundles ASR, diarization, emotion detection, and streaming into one MIT-licensed package.

20.3k Python Image · Video · Audio · explained
15
Wan-Video/Wan2.2
+12 ★/daysteady

Wan2.2 is an open video generation suite that treats denoising timesteps like specialist jobs, scaling from a 5B model on an RTX 4090 up to a 14B MoE flagship that demands 80GB VRAM.

17.5k Python Image · Video · Audio · explained
16
jianchang512/pyvideotrans
+11 ★/daysteady

pyVideoTrans wires together speech recognition, LLM translation, and voice synthesis into a single pipeline for local or API-driven video localization.

19k Python Image · Video · Audio · explained
18
xinntao/Real-ESRGAN
+10 ★/daysteady

Real-ESRGAN turns the ESRGAN research model into a practical tool for upscaling and restoring real-world images and videos using only synthetic training data.

36.8k Python Image · Video · Audio · explained
19
chidiwilliams/buzz
+19 ★/daysteady

Buzz wraps OpenAI's Whisper in a cross-platform GUI that keeps your audio data local and adds features the raw model doesn't have.

21.4k Python Image · Video · Audio · explained
20
m-bain/whisperX
+14 ★/daysteady

OpenAI's Whisper is accurate but slow and timestamp-imprecise; WhisperX bolts on batching, forced phoneme alignment, and speaker diarization to fix that.

24k Python Image · Video · Audio · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.