It bundles the entire AI drama workflow—script, storyboard, voice, and final cut—into a single pipeline you can host yourself.
Image · Video · Audio
underdogs · picking up speedVocalinux is a fully offline, GPLv3 voice dictation app that pipes transcribed text into any Linux application on X11 or Wayland.
A self-hostable infinite canvas that wires AI image generation, reference editing, and chat into one collaborative workspace.
An on-device voice toolkit that ditches the 30-second window and redundant re-processing that makes Whisper feel sluggish for live speech.
A Tauri desktop app that auto-detects a dozen local AI backends so you don't have to wrestle with Docker or API keys.
GPA aims to unify speech recognition, text-to-speech, and voice conversion in one compact autoregressive model so you can stop juggling separate audio pipelines.
It exists because cloud image generators deserve a local memory layer, a branching canvas, and a UI outside the chat thread.
Palmier Pro exists to turn a Swift-native video editor into a shared workspace where AI agents can read and write the timeline via MCP.
Most zero-shot TTS tools still demand reference transcripts or stumble across languages; this one claims to do neither.
PrintFilm corrals AI text-to-video chaos into a four-phase production pipeline for motion comics and short dramas.
Because finding the right LTX-2 checkpoint, quantization, or LoRA across Hugging Face and ComfyUI nodes is a part-time job.
It wraps open-source image-to-3D models in a desktop app so your snapshots never leave your GPU.
A polished front-end for image generation APIs that exists because your prompt history shouldn’t live in someone else’s database.
MiniMind-O packs listen-see-speak intelligence into a 0.1B-parameter model you can retrain from the first line of code on a single desktop GPU.
LocalMiniDrama wires your API keys into a Vue+Electron pipeline that turns story outlines into short-form AI video without shipping data to anyone else's cloud.
PixelRAG renders documents into screenshot tiles and retrieves them visually, preserving tables and layout that HTML parsers strip away.
FunClip exists so you can edit video by copy-pasting text instead of scrubbing timelines.
It curates GPT Image 2 prompts and packages them as copy-paste examples, an agent skill, and a lightweight CLI.
A community knowledge base that reverse-engineers hundreds of GPT-Image2 examples into structured, agent-ready prompt protocols.
MLX-VLM crams speculative decoding, continuous batching, and KV cache quantization into a Mac-native toolkit for running multimodal models locally.
