A self-hostable infinite canvas that wires AI image generation, reference editing, and chat into one collaborative workspace.
Image · Video · Audio
underdogs breaking outIt gives Claude Code and Codex a motion-design vocabulary—106 shot recipes, 161 previews, and a Remotion template—so they can direct cinematic product videos instead of generic slideshows.
Because standard Whisper sanitizes speech, and verbatim use cases actually need every filler, stutter, and false start.
It automates the full pipeline from a one-line topic to a finished Vox-style paper-collage explainer video, because coding agents should not need an editing suite to publish a film.
A privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.
It bundles the entire AI drama workflow—script, storyboard, voice, and final cut—into a single pipeline you can host yourself.
Agnes AI is a hosted multimodal API that speaks OpenAI's protocol, letting you reroute existing clients to its text, image, video, and agent models by changing a base URL.
GPA aims to unify speech recognition, text-to-speech, and voice conversion in one compact autoregressive model so you can stop juggling separate audio pipelines.
TinyEngram open-sources experiments showing that DeepSeek's Engram architecture can inject domain knowledge into Qwen and Stable Diffusion more efficiently than LoRA, with less catastrophic forgetting.
LocalMiniDrama wires your API keys into a Vue+Electron pipeline that turns story outlines into short-form AI video without shipping data to anyone else's cloud.
It exists to rebuild objects from reference photos as token-efficient, animation-ready Three.js code, using agent vision to gate quality before every sculpting pass.
A Java-based platform that pipelines LLMs, image generators, and video models into short-form drama production — script to storyboard to rendered clip.
It gives modern audio models a shared native runtime so you can stop managing Python package conflicts and start generating speech, music, and transcripts locally.
An agent skill that forces you to approve the visual metaphor and static layout before it spends your Gemini credits on video generation.
A community knowledge base that reverse-engineers hundreds of GPT-Image2 examples into structured, agent-ready prompt protocols.
Nativ bundles an embedded mlx-vlm server into a SwiftUI app to turn your Apple Silicon Mac into a local AI workspace with OpenAI-compatible APIs.
A modular agent OS that directs Seedance 2.0 video generation with film-production discipline—shot contracts, continuity rules, and retake budgets—instead of vague prompt dumps.
It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.
PixelRAG renders documents into screenshot tiles and retrieves them visually, preserving tables and layout that HTML parsers strip away.
Open-source logo maker built on Flux Pro 1.1 with a clear path from prompt to PNG.



