poloclub/diffusiondb
A large-scale dataset of 14 million images generated by Stable Diffusion with real user prompts and hyperparameters for AI research.

Not currently ranked — collecting fresh signals.
star history
DiffusionDB is a dataset containing 14 million text-to-image pairs generated by Stable Diffusion from prompts and hyperparameters specified by real users. It is designed to support research on understanding the interplay between prompts and generative models, detecting deepfakes, and designing human-AI interaction tools. The dataset is available in two subsets (2M and 14M images) on Hugging Face with accompanying metadata.
Frequently asked
- What is poloclub/diffusiondb?
- A large-scale dataset of 14 million images generated by Stable Diffusion with real user prompts and hyperparameters for AI research.
- Is diffusiondb open source?
- Yes — poloclub/diffusiondb is open source, released under the MIT license.
- What language is diffusiondb written in?
- poloclub/diffusiondb is primarily written in Python.
- How popular is diffusiondb?
- poloclub/diffusiondb has 1.4k stars on GitHub.
- Where can I find diffusiondb?
- poloclub/diffusiondb is on GitHub at https://github.com/poloclub/diffusiondb.