GigaAM exists because Russian call centers, music, and atypical speech deserve a dedicated open-source foundation model instead of hand-me-down multilingual checkpoints.
Image · Video · Audio
underdogs · picking up speedIt gives modern audio models a shared native runtime so you can stop managing Python package conflicts and start generating speech, music, and transcripts locally.
It provides the theoretically correct causal initialization that autoregressive video distillation was missing, enabling one-step to four-step generation without extra training overhead.
AICON wires script parsing, character-aware generation, and video publishing into one node-based canvas so you don't lose your actors between scenes.
It transforms scripts into consistent, AI-generated video sequences so teams don't have to manually prompt-engineer every scene.
OpenWhispr is the open-source, privacy-first alternative to WisprFlow and Granola that lets you choose between local Whisper/Parakeet models or your own cloud API keys.
It wires Alibaba's open-source Qwen3-TTS into ComfyUI so you can clone, design, and script voices by dragging nodes instead of writing Python.
A curated anthology of the most effective—and verbose—GPT Image 2 prompts scraped from X.
TypeWhisper is a native macOS app that transcribes speech using local AI models by default, then lets you chain the text through programmable workflows, cloud LLMs, or automation APIs.
A training and inference stack that squeezes autoregressive video models onto FP4 weights without making them unwatchable.
An open-source answer to paywalled AI camera-angle tools, wrapping Qwen-Image-Edit-Plus in a Three.js control rig so anyone can generate multi-angle images without a subscription.
It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.
YuE is an open foundation model that turns lyrics into full, multi-minute songs with vocals and accompaniment, offering an open-weight alternative to closed commercial generators.
Moshi is a speech-text foundation model built for real-time, full-duplex conversation, using a custom streaming neural codec that compresses 24 kHz audio to 1.1 kbps.
It turns Suno’s web app into a programmable API by automating a headless browser and outsourcing CAPTCHA solving to a paid service.
An open-source macOS alternative to Granola and WisprFlow that keeps all speech-to-text processing on Apple Silicon.
It curates GPT Image 2 prompts and packages them as copy-paste examples, an agent skill, and a lightweight CLI.
A privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.
ClipForge generates shoppable short-form videos from a single product image while automatically enforcing platform compliance rules for Chinese social commerce.
A minimal VLM you can train from scratch on one GPU in two hours for the price of a coffee.

