An open-source system that turns AI coding assistants into autonomous video studios, handling research, scripting, asset generation, and final render.
Image · Video · Audio
big names on the moveVoicebox is a local-first alternative to ElevenLabs and WisprFlow that clones voices, dictates anywhere, and gives AI agents a mouthpiece — all without shipping audio to the cloud.
It exists because clicking 'generate' isn't enough when you need to control every model, parameter, and preprocessing step.
An Electron app that wraps 200+ generative models behind a single UI, with an unusual pitch: no guardrails, no cloud lock-in, and a split personality between local and remote inference.
An offline, cross-platform dictation app that aims to be the most forkable speech-to-text tool, not the most polished one.
VoxCPM2 proves TTS doesn't need discrete tokens: a 2B-parameter diffusion model generates continuous 48kHz speech for 30 languages and text-prompted voice cloning.
A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.
VibeVoice is a family of open-source speech models from Microsoft built to ingest, transcribe, and generate very long audio sessions—up to an hour—in a single pass.
To give developers a single, general-purpose speech model that handles transcription, translation, and language identification by treating tasks as tokens to predict.
Pixelle-Video exists because producing a short video still requires scripting, generating assets, narrating, and editing; it automates all of that behind a single Web UI by orchestrating external AI services.
SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.
Deep-Live-Cam exists to turn a single selfie into a real-time webcam deepfake or video face swap running entirely on local hardware.
RVC exists because voice conversion usually demands heavy compute and hours of data; it packs training, inference, and vocal separation into a browser UI that claims to work with ten minutes of audio and weaker GPUs.
LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.
A TTS project built to clone realistic voices from just one minute of training audio.
A reimplementation of OpenAI's Whisper that trades the original inference engine for CTranslate2 and gains up to 4× speed without sacrificing accuracy.
CosyVoice provides open-source training, inference, and deployment tools for zero-shot multilingual speech synthesis using large language models.
A Tencent research project that restores degraded faces by tapping into the rich priors locked inside a pretrained StyleGAN2 model.
OpenAI's Whisper is accurate but slow and timestamp-imprecise; WhisperX bolts on batching, forced phoneme alignment, and speaker diarization to fix that.
It renders high-quality novel views of real-world scenes at 30 fps by replacing costly neural radiance fields with optimized 3D Gaussians.


