gligen/GLIGEN
A text-to-image diffusion model that grounds generation on spatial inputs like bounding boxes, keypoints, and reference images.

Not currently ranked — collecting fresh signals.
star history
GLIGEN extends frozen text-to-image models to accept additional spatial conditioning inputs including bounding boxes, keypoints, and reference images. Published at CVPR 2023, it demonstrates zero-shot performance on COCO and LVIS benchmarks that exceeds supervised layout-to-image baselines. The project includes inference code and integration with Hugging Face Spaces for demos.
Frequently asked
- What is gligen/GLIGEN?
- A text-to-image diffusion model that grounds generation on spatial inputs like bounding boxes, keypoints, and reference images.
- Is GLIGEN open source?
- Yes — gligen/GLIGEN is open source, released under the MIT license.
- What language is GLIGEN written in?
- gligen/GLIGEN is primarily written in Python.
- How popular is GLIGEN?
- gligen/GLIGEN has 2.2k stars on GitHub.
- Where can I find GLIGEN?
- gligen/GLIGEN is on GitHub at https://github.com/gligen/GLIGEN.