← all repositories
AlibabaResearch/SparkDiffusion

Alibaba squeezes DiT video generation down to a handful of steps

SparkDiffusion stacks sparse attention, low-rank approximation, and distillation to claim 200×+ faster video inference on Wan models.

SparkDiffusion
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

What it does

SparkDiffusion is an acceleration framework for Diffusion Transformer video models, targeting Wan 2.1 and Wan 2.2 (text-to-video and image-to-video). It attacks inference cost from three directions at once: sparse low-rank attention (RoLa), few-step distillation (CrossDistill), and custom Triton-based operators. The released checkpoints run in 3–4 sampling steps at 90–97% attention sparsity, and the repo claims 200×+ end-to-end speedup over the dense multi-step baseline. Weights are on Hugging Face; the pipeline covers sparse finetuning, distillation, and inference end to end.

The interesting bit

The trick is that sparsity isn’t just an inference-time hack — RoLa is used during finetuning and distillation too, so the model is actually trained to tolerate 90–97% of its attention being thrown away. The sparse-attention registry is pluggable, so new attention variants slot in without touching the training or inference core.

Key highlights

  • Joint sparsity + low-rank + distillation pipeline, with optional weight-activation quantization (w8a8) and fp8 support on capable GPUs.
  • Self-developed operators under sparkdiffusion/ops/ — no external sparse-attention checkout needed at inference time.
  • Checkpoints for Wan 2.1/2.2 T2V and I2V, from 1.3B to 14B, at 480p and 720p.
  • Sensible checkpoint loader: fails loudly on zero backbone matches, warns on partial matches, instead of silently training garbage.
  • Dense, sparse, and distilled checkpoints all run through the same inference wrappers for fair comparison.

Caveats

  • Training-side RoLa still requires an external SLA checkout (SLA_SRC), which the launchers validate before starting — a rough edge on the finetuning path.
  • Linux + CUDA GPU territory; the 265× headline figure is the best case, with 200×+ quoted as the general end-to-end number.
  • The 265× claim comes from the repo’s own description; the README itself doesn’t show the benchmark methodology.

Verdict

Worth a look if you’re running or building on Wan video models and inference cost is your bottleneck — the released checkpoints make it usable without retraining. If you’re not in the DiT video world, it’s a well-organized case study in stacking compression techniques rather than a general tool.

Frequently asked

What is AlibabaResearch/SparkDiffusion?
SparkDiffusion stacks sparse attention, low-rank approximation, and distillation to claim 200×+ faster video inference on Wan models.
Is SparkDiffusion open source?
Yes — AlibabaResearch/SparkDiffusion is open source, released under the Apache-2.0 license.
What language is SparkDiffusion written in?
AlibabaResearch/SparkDiffusion is primarily written in Python.
How popular is SparkDiffusion?
AlibabaResearch/SparkDiffusion has 533 stars on GitHub.
Where can I find SparkDiffusion?
AlibabaResearch/SparkDiffusion is on GitHub at https://github.com/AlibabaResearch/SparkDiffusion.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.