songweige/rich-text-to-image
A diffusion model system that uses rich text formatting—font size, color, style, footnotes—to control text-to-image generation.

Not currently ranked — collecting fresh signals.
star history
This research project enables fine-grained control over AI image generation by leveraging formatting information from rich text documents. It extends Stable Diffusion and SD-XL with capabilities for explicit token reweighting, precise color rendering, local style control, and detailed region synthesis. The project includes a HuggingFace demo and an Automatic1111 WebUI extension for practical use.
Frequently asked
- What is songweige/rich-text-to-image?
- A diffusion model system that uses rich text formatting—font size, color, style, footnotes—to control text-to-image generation.
- Is rich-text-to-image open source?
- Yes — songweige/rich-text-to-image is open source, released under the MIT license.
- What language is rich-text-to-image written in?
- songweige/rich-text-to-image is primarily written in Python.
- How popular is rich-text-to-image?
- songweige/rich-text-to-image has 800 stars on GitHub.
- Where can I find rich-text-to-image?
- songweige/rich-text-to-image is on GitHub at https://github.com/songweige/rich-text-to-image.