← all repositories

FoundationVision/UniTok

A unified visual tokenizer that converts images into discrete tokens for use in autoregressive generation and multimodal understanding models.

UniTok
Not currently ranked — collecting fresh signals.
star history

UniTok is a unified visual tokenizer designed for both visual generation and understanding tasks. It provides discrete tokenization of images compatible with autoregressive generative models like LlamaGen and multimodal understanding models like LLaVA. The tokenizer supports unified multimodal LLMs including Chameleon and Liquid, enabling both image generation and comprehension within the same framework. It was published at NeurIPS 2025 as a Spotlight paper.

Frequently asked

What is FoundationVision/UniTok?
A unified visual tokenizer that converts images into discrete tokens for use in autoregressive generation and multimodal understanding models.
Is UniTok open source?
Yes — FoundationVision/UniTok is open source, released under the MIT license.
What language is UniTok written in?
FoundationVision/UniTok is primarily written in Python.
How popular is UniTok?
FoundationVision/UniTok has 529 stars on GitHub.
Where can I find UniTok?
FoundationVision/UniTok is on GitHub at https://github.com/FoundationVision/UniTok.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.