Image · Video · Audio

Image · Video · Audio

big names · picking up speed
02
calesthio/OpenMontage
+354 ★/dayaccelerating

An open-source system that turns AI coding assistants into autonomous video studios, handling research, scripting, asset generation, and final render.

47.8k Python Agents · explained Feature
03
index-tts/index-tts
+46 ★/dayaccelerating

IndexTTS2 is a zero-shot text-to-speech system that disentangles speaker timbre from emotional expression so you can clone a voice and dial in a mood separately.

22.8k Python Image · Video · Audio · explained
04
ggml-org/whisper.cpp
+36 ★/dayaccelerating

A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.

52.9k C++ Image · Video · Audio · explained
05
Anil-matcha/Open-Generative-AI
+77 ★/dayaccelerating

An Electron app that wraps 200+ generative models behind a single UI, with an unusual pitch: no guardrails, no cloud lock-in, and a split personality between local and remote inference.

26.2k JavaScript Image · Video · Audio · explained
06
QwenAudio/CosyVoice
+18 ★/dayaccelerating

CosyVoice provides open-source training, inference, and deployment tools for zero-shot multilingual speech synthesis using large language models.

22.7k Python Image · Video · Audio · explained
07
OpenBMB/MiniCPM-V
+8.6 ★/dayaccelerating

To squeeze multimodal understanding—vision, video, and even real-time speech—into models small enough to run natively on a handset.

26.2k Python Image · Video · Audio · explained
08
graphdeco-inria/gaussian-splatting
+9.0 ★/dayaccelerating

It renders high-quality novel views of real-world scenes at 30 fps by replacing costly neural radiance fields with optimized 3D Gaussians.

22.9k Python Computer Vision · explained
11
microsoft/VibeVoice
+91 ★/daysteady

VibeVoice is a family of open-source speech models from Microsoft built to ingest, transcribe, and generate very long audio sessions—up to an hour—in a single pass.

52.6k Python Image · Video · Audio · explained
13
invoke-ai/InvokeAI
+13 ★/daysteady

It gives artists and professionals a local, node-based studio for Stable Diffusion, treating AI as a collaborator rather than an opaque generator.

27.9k Python Image · Video · Audio · explained
14
m-bain/whisperX
+15 ★/daycooling

OpenAI's Whisper is accurate but slow and timestamp-imprecise; WhisperX bolts on batching, forced phoneme alignment, and speaker diarization to fix that.

23.5k Python Image · Video · Audio · explained
15
myshell-ai/OpenVoice
+5.9 ★/daycooling

OpenVoice clones a speaker's tone color and lets you reshape emotion, accent, and language independently—even generating speech in languages absent from the training data.

37.1k Python Image · Video · Audio · explained
16
2noise/ChatTTS
+4.9 ★/daycooling

ChatTTS is a generative speech model built for dialogue, letting LLM assistants laugh, pause, and speak in multiple voices.

39.8k Python Image · Video · Audio · explained
17
openai/CLIP
+3.9 ★/daycooling

CLIP learns shared image-text representations so you can label photos with natural language instead of curated datasets.

34.2k Jupyter Notebook Language Models · explained
18
TencentARC/GFPGAN
+0.1 ★/daycooling

A Tencent research project that restores degraded faces by tapping into the rich priors locked inside a pretrained StyleGAN2 model.

37.7k Python Computer Vision · explained
19
xinntao/Real-ESRGAN
+9.3 ★/daycooling

Real-ESRGAN turns the ESRGAN research model into a practical tool for upscaling and restoring real-world images and videos using only synthetic training data.

36.5k Python Image · Video · Audio · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.