jingyi0000/VLM_survey
A systematic survey of vision-language models applied to visual recognition tasks including classification, detection, and segmentation.

Not currently ranked — collecting fresh signals.
star history
This repository hosts a comprehensive survey of Vision-Language Models (VLMs) compiled as an academic resource. It catalogs VLM studies across various visual recognition tasks such as image classification, object detection, and semantic segmentation. The survey, published in IEEE TPAMI 2024, serves as an curated awesome list of research papers in the multi-modal/VLM space.
Frequently asked
- What is jingyi0000/VLM_survey?
- A systematic survey of vision-language models applied to visual recognition tasks including classification, detection, and segmentation.
- Is VLM_survey open source?
- Yes — jingyi0000/VLM_survey is an open-source project tracked on heatdrop.
- How popular is VLM_survey?
- jingyi0000/VLM_survey has 3.1k stars on GitHub.
- Where can I find VLM_survey?
- jingyi0000/VLM_survey is on GitHub at https://github.com/jingyi0000/VLM_survey.