ByteVisionLab/TokenFlow
Unified image tokenizer with dual-codebook architecture for multimodal understanding and text-to-image generation.

Not currently ranked — collecting fresh signals.
star history
TokenFlow bridges multimodal understanding and generation using an innovative dual-codebook architecture that decouples semantic and pixel-level feature learning while maintaining alignment through a shared mapping mechanism. It surpasses flagship models like LLaVA-1.5 and EMU3 on visual comprehension tasks and achieves performance comparable to SDXL on text-to-image generation at 256×256 resolution.
Frequently asked
- What is ByteVisionLab/TokenFlow?
- Unified image tokenizer with dual-codebook architecture for multimodal understanding and text-to-image generation.
- Is TokenFlow open source?
- Yes — ByteVisionLab/TokenFlow is open source, released under the Apache-2.0 license.
- What language is TokenFlow written in?
- ByteVisionLab/TokenFlow is primarily written in Python.
- How popular is TokenFlow?
- ByteVisionLab/TokenFlow has 464 stars on GitHub.
- Where can I find TokenFlow?
- ByteVisionLab/TokenFlow is on GitHub at https://github.com/ByteVisionLab/TokenFlow.