Image · Video · Audio

Image · Video · Audio

big names on the move
01
calesthio/OpenMontage
+792 ★/dayaccelerating

An open-source system that turns AI coding assistants into autonomous video studios, handling research, scripting, asset generation, and final render.

42.3k Python Agents · explained Feature
02
jamiepine/voicebox
+565 ★/daycooling

Voicebox is a local-first alternative to ElevenLabs and WisprFlow that clones voices, dictates anywhere, and gives AI agents a mouthpiece — all without shipping audio to the cloud.

46.7k TypeScript Image · Video · Audio · explained Feature
03
Comfy-Org/ComfyUI
+131 ★/daycooling

It exists because clicking 'generate' isn't enough when you need to control every model, parameter, and preprocessing step.

122.2k Python Image · Video · Audio · explained
04
Anil-matcha/Open-Generative-AI
+113 ★/daycooling

An Electron app that wraps 200+ generative models behind a single UI, with an unusual pitch: no guardrails, no cloud lock-in, and a split personality between local and remote inference.

24.8k JavaScript Image · Video · Audio · explained
05
cjpais/Handy
+97 ★/dayaccelerating

An offline, cross-platform dictation app that aims to be the most forkable speech-to-text tool, not the most polished one.

27.5k Rust Image · Video · Audio · explained
06
OpenBMB/VoxCPM
+75 ★/daycooling

VoxCPM2 proves TTS doesn't need discrete tokens: a 2B-parameter diffusion model generates continuous 48kHz speech for 30 languages and text-prompted voice cloning.

34.2k Python Image · Video · Audio · explained
07
ggml-org/whisper.cpp
+64 ★/dayaccelerating

A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.

52.3k C++ Image · Video · Audio · explained
08
microsoft/VibeVoice
+53 ★/dayaccelerating

VibeVoice is a family of open-source speech models from Microsoft built to ingest, transcribe, and generate very long audio sessions—up to an hour—in a single pass.

50.5k Python Image · Video · Audio · explained
09
openai/whisper
+52 ★/daycooling

To give developers a single, general-purpose speech model that handles transcription, translation, and language identification by treating tasks as tokens to predict.

105.6k Python Image · Video · Audio · explained
10
ATH-MaaS/Pixelle-Video
+43 ★/daycooling

Pixelle-Video exists because producing a short video still requires scripting, generating assets, narrating, and editing; it automates all of that behind a single Web UI by orchestrating external AI services.

26k Python Image · Video · Audio · explained
11
sgl-project/sglang
+40 ★/dayaccelerating

SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.

30.7k Python Inference · Serving · explained
12
hacksider/Deep-Live-Cam
+37 ★/dayaccelerating

Deep-Live-Cam exists to turn a single selfie into a real-time webcam deepfake or video face swap running entirely on local hardware.

95.2k Python Image · Video · Audio · explained
14
mudler/LocalAI
+28 ★/dayaccelerating

LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.

47.9k Go Inference · Serving · explained
16
SYSTRAN/faster-whisper
+24 ★/dayaccelerating

A reimplementation of OpenAI's Whisper that trades the original inference engine for CTranslate2 and gains up to 4× speed without sacrificing accuracy.

24.5k Python Inference · Serving · explained
17
QwenAudio/CosyVoice
+22 ★/daysteady

CosyVoice provides open-source training, inference, and deployment tools for zero-shot multilingual speech synthesis using large language models.

22.4k Python Image · Video · Audio · explained
18
TencentARC/GFPGAN
+18 ★/dayaccelerating

A Tencent research project that restores degraded faces by tapping into the rich priors locked inside a pretrained StyleGAN2 model.

37.6k Python Computer Vision · explained
19
m-bain/whisperX
+17 ★/daysteady

OpenAI's Whisper is accurate but slow and timestamp-imprecise; WhisperX bolts on batching, forced phoneme alignment, and speaker diarization to fix that.

23.3k Python Image · Video · Audio · explained
20
graphdeco-inria/gaussian-splatting
+13 ★/dayaccelerating

It renders high-quality novel views of real-world scenes at 30 fps by replacing costly neural radiance fields with optimized 3D Gaussians.

22.8k Python Computer Vision · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.