chengzeyi/stable-fast
A high-performance inference optimization framework for diffusion models using PyTorch, CUDA, and OpenAI Triton on NVIDIA GPUs.

Not currently ranked — collecting fresh signals.
star history
Stable Fast is an inference acceleration framework that achieves state-of-the-art performance for all HuggingFace Diffusers pipelines including the latest Stable Video Diffusion. Unlike TensorRT or AITemplate which require lengthy compilation, it compiles models in seconds. It natively supports dynamic shapes, LoRA adapters, and ControlNet while targeting NVIDIA GPUs.
Frequently asked
- What is chengzeyi/stable-fast?
- A high-performance inference optimization framework for diffusion models using PyTorch, CUDA, and OpenAI Triton on NVIDIA GPUs.
- Is stable-fast open source?
- Yes — chengzeyi/stable-fast is open source, released under the MIT license.
- What language is stable-fast written in?
- chengzeyi/stable-fast is primarily written in Python.
- How popular is stable-fast?
- chengzeyi/stable-fast has 1.3k stars on GitHub.
- Where can I find stable-fast?
- chengzeyi/stable-fast is on GitHub at https://github.com/chengzeyi/stable-fast.