OmniVoice Studio bundles voice cloning, dubbing, dictation, and TTS into a desktop app that keeps all audio processing off the internet and away from API keys.
Image · Video · Audio
big names on the moveA community knowledge base that reverse-engineers hundreds of GPT-Image2 examples into structured, agent-ready prompt protocols.
SGLang exists to push low-latency, high-throughput inference for LLMs and multimodal models from a single GPU up to massive clusters.
An open-source system that turns AI coding assistants into autonomous video studios, handling research, scripting, asset generation, and final render.
It exists because clicking 'generate' isn't enough when you need to control every model, parameter, and preprocessing step.
Voicebox is a local-first alternative to ElevenLabs and WisprFlow that clones voices, dictates anywhere, and gives AI agents a mouthpiece — all without shipping audio to the cloud.
An Electron app that wraps 200+ generative models behind a single UI, with an unusual pitch: no guardrails, no cloud lock-in, and a split personality between local and remote inference.
To give developers a single, general-purpose speech model that handles transcription, translation, and language identification by treating tasks as tokens to predict.
VibeVoice is a family of open-source speech models from Microsoft built to ingest, transcribe, and generate very long audio sessions—up to an hour—in a single pass.
An offline, cross-platform dictation app that aims to be the most forkable speech-to-text tool, not the most polished one.
VoxCPM2 proves TTS doesn't need discrete tokens: a 2B-parameter diffusion model generates continuous 48kHz speech for 30 languages and text-prompted voice cloning.
Pixelle-Video exists because producing a short video still requires scripting, generating assets, narrating, and editing; it automates all of that behind a single Web UI by orchestrating external AI services.
LocalAI wraps 36+ inference engines behind one OpenAI-compatible API and pulls them on demand, so you can run LLMs, vision, voice, and video on anything from a CPU to a Jetson.
A TTS project built to clone realistic voices from just one minute of training audio.
A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.
IndexTTS2 is a zero-shot text-to-speech system that disentangles speaker timbre from emotional expression so you can clone a voice and dial in a mood separately.
Deezer open-sourced its TensorFlow stem splitter so developers can pull vocals, drums, bass, and piano out of a mixed track without training a model from scratch.
RVC exists because voice conversion usually demands heavy compute and hours of data; it packs training, inference, and vocal separation into a browser UI that claims to work with ten minutes of audio and weaker GPUs.
Buzz wraps OpenAI's Whisper in a cross-platform GUI that keeps your audio data local and adds features the raw model doesn't have.
Deep-Live-Cam exists to turn a single selfie into a real-time webcam deepfake or video face swap running entirely on local hardware.


