← all repositories

baaivision/tokenize-anything

A promptable vision foundation model for segmenting, recognizing, and captioning arbitrary regions with flexible visual prompts.

601 stars Jupyter Notebook Image · Video · AudioLanguage Models
tokenize-anything
Not currently ranked — collecting fresh signals.
star history

Tokenize Anything via Prompting (TAP) is a unified model that simultaneously performs open-world segmentation, recognition, and captioning using visual prompts such as points, boxes, and sketches. The model is trained on exhaustive segmentation masks from SA-1B combined with semantic priors from EVA-CLIP, a 5-billion parameter vision-language model. It provides a modular design with decoupled components and predictors for flexible integration.

Frequently asked

What is baaivision/tokenize-anything?
A promptable vision foundation model for segmenting, recognizing, and captioning arbitrary regions with flexible visual prompts.
Is tokenize-anything open source?
Yes — baaivision/tokenize-anything is open source, released under the Apache-2.0 license.
What language is tokenize-anything written in?
baaivision/tokenize-anything is primarily written in Jupyter Notebook.
How popular is tokenize-anything?
baaivision/tokenize-anything has 601 stars on GitHub.
Where can I find tokenize-anything?
baaivision/tokenize-anything is on GitHub at https://github.com/baaivision/tokenize-anything.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.