open-mmlab/Multimodal-GPT
Multi-modal chatbot that processes visual and language instructions, based on the OpenFlamingo vision-language model.

Not currently ranked — collecting fresh signals.
star history
This repository trains a multi-modal chatbot by fine-tuning the OpenFlamingo architecture on both visual and language instruction datasets. It creates training data across VQA, image captioning, visual reasoning, text OCR, and visual dialogue tasks. The project performs joint training of visual and language instructions to improve model performance on multi-modal instruction following.
Frequently asked
- What is open-mmlab/Multimodal-GPT?
- Multi-modal chatbot that processes visual and language instructions, based on the OpenFlamingo vision-language model.
- Is Multimodal-GPT open source?
- Yes — open-mmlab/Multimodal-GPT is open source, released under the Apache-2.0 license.
- What language is Multimodal-GPT written in?
- open-mmlab/Multimodal-GPT is primarily written in Python.
- How popular is Multimodal-GPT?
- open-mmlab/Multimodal-GPT has 1.5k stars on GitHub.
- Where can I find Multimodal-GPT?
- open-mmlab/Multimodal-GPT is on GitHub at https://github.com/open-mmlab/Multimodal-GPT.