TencentQQGYLab/ELLA
ELLA equips diffusion models with large language models to improve semantic alignment in text-to-image generation.

Not currently ranked — collecting fresh signals.
star history
ELLA is a research project that combines diffusion models with LLMs to enhance semantic alignment in image generation. The approach allows text-to-image diffusion models to better understand and follow complex text prompts by integrating large language model capabilities. The repository also includes EMMA, a related technique that enables text-to-image models to accept multi-modal prompts.
Frequently asked
- What is TencentQQGYLab/ELLA?
- ELLA equips diffusion models with large language models to improve semantic alignment in text-to-image generation.
- Is ELLA open source?
- Yes — TencentQQGYLab/ELLA is open source, released under the Apache-2.0 license.
- What language is ELLA written in?
- TencentQQGYLab/ELLA is primarily written in Python.
- How popular is ELLA?
- TencentQQGYLab/ELLA has 1.3k stars on GitHub.
- Where can I find ELLA?
- TencentQQGYLab/ELLA is on GitHub at https://github.com/TencentQQGYLab/ELLA.