← all repositories

NVlabs/prismer

Prismer is a vision-language model that uses pre-trained experts across multiple vision-language tasks including image captioning and visual question answering.

prismer
Not currently ranked — collecting fresh signals.
star history

Prismer implements a vision-language architecture combining multiple pre-trained expert models to handle diverse vision-language tasks. The model supports image captioning, visual question answering, and other multimodal tasks through a multi-task expert framework. It is built on PyTorch with Hugging Face accelerate for distributed multi-node multi-gpu training. A demo is available via HuggingFace Spaces.

Frequently asked

What is NVlabs/prismer?
Prismer is a vision-language model that uses pre-trained experts across multiple vision-language tasks including image captioning and visual question answering.
Is prismer open source?
Yes — NVlabs/prismer is an open-source project tracked on heatdrop.
What language is prismer written in?
NVlabs/prismer is primarily written in Python.
How popular is prismer?
NVlabs/prismer has 1.3k stars on GitHub.
Where can I find prismer?
NVlabs/prismer is on GitHub at https://github.com/NVlabs/prismer.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.