A self-hostable infinite canvas that wires AI image generation, reference editing, and chat into one collaborative workspace.
Image · Video · Audio
underdogs · picking up speedA privacy-first Android fork that runs LLMs, image generation, and speech AI entirely offline, then locks itself behind your fingerprint.
It turns any LLM into a ComfyUI operator that edits live graphs, manages models, and runs workflows instead of just forwarding prompts.
Agnes AI is a hosted multimodal API that speaks OpenAI's protocol, letting you reroute existing clients to its text, image, video, and agent models by changing a base URL.
A Japanese TTS system that lets you control speaking style by sprinkling emoji into the input text.
A community knowledge base that reverse-engineers hundreds of GPT-Image2 examples into structured, agent-ready prompt protocols.
It plugs nine image and video models into your AI coding tools via MCP, letting them generate visuals in parallel without cluttering your chat context.
A grab-bag node pack whose Set/Get rewrite might finally tame your worst workflow tangles.
Most AI image prompts are one-off text blobs; this repo distills 96 visual styles into structured JSON templates so you can swap variables without losing style direction.
An inference framework that corrals a dozen video and image generation models into a single runtime optimized to outrun Diffusers and FastVideo on both H100s and RTX 4090Ds.
It unifies Stable Diffusion, GGUF chat, Whisper, and Kokoro TTS into a single offline desktop GUI so you can skip cloud APIs, subscriptions, and censorship filters.
It gives Claude Code and Codex a motion-design vocabulary—106 shot recipes, 161 previews, and a Remotion template—so they can direct cinematic product videos instead of generic slideshows.
It curates GPT Image 2 prompts and packages them as copy-paste examples, an agent skill, and a lightweight CLI.
Applio offers a browser-based voice conversion toolkit that prioritizes stability and plugin extensibility over constant feature churn.
It wants to be the only browser tab you need for local AI image and video generation, serving both beginners and node-graph tinkerers.
It gives you polyphonic audio-to-MIDI transcription, complete with pitch bends, without the resource bill of heavier research tools.
It generates lip-synced faces by feeding Whisper audio embeddings straight into a latent diffusion U-Net, skipping the usual motion-representation detour.
It wraps open-source image-to-3D models in a desktop app so your snapshots never leave your GPU.
SD.Next exists to run Stable Diffusion and related models on virtually any consumer hardware while bundling quantization, captioning, and video tools into one interface.
A polished front-end for image generation APIs that exists because your prompt history shouldn’t live in someone else’s database.





