← all repositories

open-mmlab/Multimodal-GPT

Multi-modal chatbot that processes visual and language instructions, based on the OpenFlamingo vision-language model.

Multimodal-GPT
Not currently ranked — collecting fresh signals.
star history

This repository trains a multi-modal chatbot by fine-tuning the OpenFlamingo architecture on both visual and language instruction datasets. It creates training data across VQA, image captioning, visual reasoning, text OCR, and visual dialogue tasks. The project performs joint training of visual and language instructions to improve model performance on multi-modal instruction following.

Frequently asked

What is open-mmlab/Multimodal-GPT?
Multi-modal chatbot that processes visual and language instructions, based on the OpenFlamingo vision-language model.
Is Multimodal-GPT open source?
Yes — open-mmlab/Multimodal-GPT is open source, released under the Apache-2.0 license.
What language is Multimodal-GPT written in?
open-mmlab/Multimodal-GPT is primarily written in Python.
How popular is Multimodal-GPT?
open-mmlab/Multimodal-GPT has 1.5k stars on GitHub.
Where can I find Multimodal-GPT?
open-mmlab/Multimodal-GPT is on GitHub at https://github.com/open-mmlab/Multimodal-GPT.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.