← all repositories

THU-SI/Spatial-MLLM

Spatial-MLLM enhances existing video multimodal LLMs with visual-based spatial intelligence capabilities.

472 stars Python Language Models
Spatial-MLLM
Not currently ranked — collecting fresh signals.
star history

Spatial-MLLM is a method that significantly enhances the visual-based spatial intelligence of existing video multimodal large language models. The project provides supervised fine-tuning training code, evaluation code, and pre-trained models for spatial reasoning tasks. It achieves state-of-the-art performance on benchmarks like VSI-Bench and releases models trained on datasets such as Spatial-MLLM-120k.

Frequently asked

What is THU-SI/Spatial-MLLM?
Spatial-MLLM enhances existing video multimodal LLMs with visual-based spatial intelligence capabilities.
Is Spatial-MLLM open source?
Yes — THU-SI/Spatial-MLLM is open source, released under the MIT license.
What language is Spatial-MLLM written in?
THU-SI/Spatial-MLLM is primarily written in Python.
How popular is Spatial-MLLM?
THU-SI/Spatial-MLLM has 472 stars on GitHub.
Where can I find Spatial-MLLM?
THU-SI/Spatial-MLLM is on GitHub at https://github.com/THU-SI/Spatial-MLLM.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.