Image · Video · Audio

Image · Video · Audio

underdogs · picking up speed
01
dramaclaw/dramaclaw
+69% /wk +205 ★/dayaccelerating

It bundles the entire AI drama workflow—script, storyboard, voice, and final cut—into a single pipeline you can host yourself.

2.1k TypeScript Image · Video · Audio · explained Feature
02
jatinkrmalik/vocalinux
+15% /wk +15 ★/dayaccelerating

Vocalinux is a fully offline, GPLv3 voice dictation app that pipes transcribed text into any Linux application on X11 or Wayland.

668 Python Image · Video · Audio · explained
03
basketikun/infinite-canvas
+20% /wk +109 ★/dayaccelerating

A self-hostable infinite canvas that wires AI image generation, reference editing, and chat into one collaborative workspace.

3.8k TypeScript Creative · Design · explained
04
moonshine-ai/moonshine
+17% /wk +252 ★/dayaccelerating

An on-device voice toolkit that ditches the 30-second window and redundant re-processing that makes Whisper feel sluggish for live speech.

10.4k C++ Image · Video · Audio · explained
06
AutoArk/GPA
+21% /wk +52 ★/dayaccelerating

GPA aims to unify speech recognition, text-to-speech, and voice conversion in one compact autoregressive model so you can stop juggling separate audio pipelines.

1.7k Python Image · Video · Audio · explained
07
lidge-jun/ima2-gen
+12% /wk +10 ★/dayaccelerating

It exists because cloud image generators deserve a local memory layer, a branching canvas, and a UI outside the chat thread.

603 TypeScript Image · Video · Audio · explained
08
palmier-io/palmier-pro
+12% /wk +208 ★/dayaccelerating

Palmier Pro exists to turn a Swift-native video editor into a shared workspace where AI agents can read and write the timeline via MCP.

12.2k Swift Agents · explained Feature
10
yuanzhongqiao/printfilm
+5.0% /wk +21 ★/dayaccelerating

PrintFilm corrals AI text-to-video chaos into a four-phase production pipeline for motion comics and short dramas.

2.9k TypeScript Image · Video · Audio · explained
11
wildminder/awesome-ltx2
+5.9% /wk +4.7 ★/dayaccelerating

Because finding the right LTX-2 checkpoint, quantization, or LoRA across Hugging Face and ComfyUI nodes is a part-time job.

555 Image · Video · Audio · explained
14
jingyaogong/minimind-o
+4.6% /wk +14 ★/dayaccelerating

MiniMind-O packs listen-see-speak intelligence into a 0.1B-parameter model you can retrain from the first line of code on a single desktop GPU.

2.2k Python Language Models · explained
15
xuanyustudio/LocalMiniDrama
+11% /wk +15 ★/dayaccelerating

LocalMiniDrama wires your API keys into a Vue+Electron pipeline that turns story outlines into short-form AI video without shipping data to anyone else's cloud.

971 JavaScript Image · Video · Audio · explained
16
StarTrail-org/PixelRAG
+6.4% /wk +66 ★/dayaccelerating

PixelRAG renders documents into screenshot tiles and retrieves them visually, preserving tables and layout that HTML parsers strip away.

7.2k Python RAG · Search · explained
17
modelscope/FunClip
+2.9% /wk +25 ★/dayaccelerating

FunClip exists so you can edit video by copy-pasting text instead of scrubbing timelines.

6.1k Python Domain Apps · explained
18
wuyoscar/GPT-Image2-Skill
+5.5% /wk +31 ★/dayaccelerating

It curates GPT Image 2 prompts and packages them as copy-paste examples, an agent skill, and a lightweight CLI.

4k Python Coding Assistants · explained
20
Blaizzy/mlx-vlm
+2.2% /wk +16 ★/dayaccelerating

MLX-VLM crams speculative decoding, continuous batching, and KV cache quantization into a Mac-native toolkit for running multimodal models locally.

5.3k Python Image · Video · Audio · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.