UMass-Embodied-AGI/3D-LLM
A large language model that accepts 3D representations (objects and scenes) as inputs for 3D visual reasoning tasks.

Not currently ranked — collecting fresh signals.
star history
3D-LLM is the first LLM system designed to process 3D point clouds and scene data from sources like ScanNet and Objaverse. The model can perform 3D captioning, question answering, and task planning on 3D environments by encoding spatial and semantic information into the language model. It uses LAVIS as its underlying vision-language framework and provides pretrained and fine-tuned checkpoints for downstream 3D understanding tasks.
Frequently asked
- What is UMass-Embodied-AGI/3D-LLM?
- A large language model that accepts 3D representations (objects and scenes) as inputs for 3D visual reasoning tasks.
- Is 3D-LLM open source?
- Yes — UMass-Embodied-AGI/3D-LLM is open source, released under the MIT license.
- What language is 3D-LLM written in?
- UMass-Embodied-AGI/3D-LLM is primarily written in Python.
- How popular is 3D-LLM?
- UMass-Embodied-AGI/3D-LLM has 1.2k stars on GitHub.
- Where can I find 3D-LLM?
- UMass-Embodied-AGI/3D-LLM is on GitHub at https://github.com/UMass-Embodied-AGI/3D-LLM.