Image · Video · Audio

Image · Video · Audio

big names on the move
01
debpalash/VoiceStudio
+1362 ★/dayaccelerating

OmniVoice Studio bundles voice cloning, dubbing, dictation, and TTS into a desktop app that keeps all audio processing off the internet and away from API keys.

22.3k Python Image · Video · Audio · explained Feature
03
sgl-project/sglang
+243 ★/dayaccelerating

SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.

35.8k Python Inference · Serving · explained
04
calesthio/OpenMontage
+153 ★/daycooling

An open-source system that turns AI coding assistants into autonomous video studios, handling research, scripting, asset generation, and final render.

57.1k Python Agents · explained Feature
06
jamiepine/voicebox
+99 ★/dayaccelerating

Voicebox is a local-first alternative to ElevenLabs and WisprFlow that clones voices, dictates anywhere, and gives AI agents a mouthpiece — all without shipping audio to the cloud.

52.9k TypeScript Image · Video · Audio · explained Feature
07
Anil-matcha/Open-Generative-AI
+83 ★/dayaccelerating

An Electron app that wraps 200+ generative models behind a single UI, with an unusual pitch: no guardrails, no cloud lock-in, and a split personality between local and remote inference.

28.3k JavaScript Image · Video · Audio · explained
08
openai/whisper
+70 ★/dayaccelerating

To give developers a single, general-purpose speech model that handles transcription, translation, and language identification by treating tasks as tokens to predict.

108.9k Python Image · Video · Audio · explained
09
microsoft/VibeVoice
+60 ★/dayaccelerating

VibeVoice is a family of open-source speech models from Microsoft built to ingest, transcribe, and generate very long audio sessions—up to an hour—in a single pass.

54k Python Image · Video · Audio · explained
10
cjpais/Handy
+54 ★/daycooling

An offline, cross-platform dictation app that aims to be the most forkable speech-to-text tool, not the most polished one.

31.3k Rust Image · Video · Audio · explained
11
OpenBMB/VoxCPM
+40 ★/daycooling

VoxCPM2 proves TTS doesn't need discrete tokens: a 2B-parameter diffusion model generates continuous 48kHz speech for 30 languages and text-prompted voice cloning.

36.9k Python Image · Video · Audio · explained
12
ATH-MaaS/Pixelle-Video
+35 ★/daycooling

Pixelle-Video exists because producing a short video still requires scripting, generating assets, narrating, and editing; it automates all of that behind a single Web UI by orchestrating external AI services.

28k Python Image · Video · Audio · explained
13
mudler/LocalAI
+28 ★/dayaccelerating

LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.

49k Go Inference · Serving · explained
15
ggml-org/whisper.cpp
+24 ★/daycooling

A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.

53.6k C++ Image · Video · Audio · explained
16
index-tts/index-tts
+23 ★/daycooling

IndexTTS2 is a zero-shot text-to-speech system that disentangles speaker timbre from emotional expression so you can clone a voice and dial in a mood separately.

23.9k Python Image · Video · Audio · explained
17
deezer/spleeter
+22 ★/dayaccelerating

Deezer open-sourced its TensorFlow stem splitter so developers can pull vocals, drums, bass, and piano out of a mixed track without training a model from scratch.

28.4k Python Image · Video · Audio · explained
19
chidiwilliams/buzz
+19 ★/daysteady

Buzz wraps OpenAI's Whisper in a cross-platform GUI that keeps your audio data local and adds features the raw model doesn't have.

21.4k Python Image · Video · Audio · explained
20
hacksider/Deep-Live-Cam
+17 ★/daycooling

Deep-Live-Cam exists to turn a single selfie into a real-time webcam deepfake or video face swap running entirely on local hardware.

96.6k Python Image · Video · Audio · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.