THU-SI/Spatial-MLLM
Spatial-MLLM enhances existing video multimodal LLMs with visual-based spatial intelligence capabilities.

Not currently ranked — collecting fresh signals.
star history
Spatial-MLLM is a method that significantly enhances the visual-based spatial intelligence of existing video multimodal large language models. The project provides supervised fine-tuning training code, evaluation code, and pre-trained models for spatial reasoning tasks. It achieves state-of-the-art performance on benchmarks like VSI-Bench and releases models trained on datasets such as Spatial-MLLM-120k.
Frequently asked
- What is THU-SI/Spatial-MLLM?
- Spatial-MLLM enhances existing video multimodal LLMs with visual-based spatial intelligence capabilities.
- Is Spatial-MLLM open source?
- Yes — THU-SI/Spatial-MLLM is open source, released under the MIT license.
- What language is Spatial-MLLM written in?
- THU-SI/Spatial-MLLM is primarily written in Python.
- How popular is Spatial-MLLM?
- THU-SI/Spatial-MLLM has 472 stars on GitHub.
- Where can I find Spatial-MLLM?
- THU-SI/Spatial-MLLM is on GitHub at https://github.com/THU-SI/Spatial-MLLM.