It exists because clicking 'generate' isn't enough when you need to control every model, parameter, and preprocessing step.
Image · Video · Audio
big names · picking up speedAn open-source system that turns AI coding assistants into autonomous video studios, handling research, scripting, asset generation, and final render.
IndexTTS2 is a zero-shot text-to-speech system that disentangles speaker timbre from emotional expression so you can clone a voice and dial in a mood separately.
A minimal C/C++ port of OpenAI’s Whisper built to transcribe speech locally on phones, browsers, and underclocked POWER9 boxes.
An Electron app that wraps 200+ generative models behind a single UI, with an unusual pitch: no guardrails, no cloud lock-in, and a split personality between local and remote inference.
CosyVoice provides open-source training, inference, and deployment tools for zero-shot multilingual speech synthesis using large language models.
To squeeze multimodal understanding—vision, video, and even real-time speech—into models small enough to run natively on a handset.
It renders high-quality novel views of real-world scenes at 30 fps by replacing costly neural radiance fields with optimized 3D Gaussians.
A TTS project built to clone realistic voices from just one minute of training audio.
Because you shouldn't need to reverse-engineer a black box just to swap a noise scheduler or fine-tune a diffusion model.
VibeVoice is a family of open-source speech models from Microsoft built to ingest, transcribe, and generate very long audio sessions—up to an hour—in a single pass.
It collects the best social-media experiments with a Gemini-2.5-flash-image derivative and releases a 150k identity-consistent dataset for the community.
It gives artists and professionals a local, node-based studio for Stable Diffusion, treating AI as a collaborator rather than an opaque generator.
OpenAI's Whisper is accurate but slow and timestamp-imprecise; WhisperX bolts on batching, forced phoneme alignment, and speaker diarization to fix that.
OpenVoice clones a speaker's tone color and lets you reshape emotion, accent, and language independently—even generating speech in languages absent from the training data.
ChatTTS is a generative speech model built for dialogue, letting LLM assistants laugh, pause, and speak in multiple voices.
CLIP learns shared image-text representations so you can label photos with natural language instead of curated datasets.
A Tencent research project that restores degraded faces by tapping into the rich priors locked inside a pretrained StyleGAN2 model.
Real-ESRGAN turns the ESRGAN research model into a practical tool for upscaling and restoring real-world images and videos using only synthetic training data.
Trained its own deep-learning models so you don't have to sing along.

