← all repositories

FoundationVision/Groma

Groma is a multimodal LLM that uses localized visual tokenization to enable region-level understanding and visual grounding capabilities.

Groma
Not currently ranked — collecting fresh signals.
star history

Groma is a grounded multimodal large language model that introduces visual tokenization for localization, allowing it to process user-defined region inputs (bounding boxes) and generate responses grounded to specific visual regions. It achieves state-of-the-art performance on referring expression comprehension benchmarks like RefCOCO, RefCOCO+, and RefCOCOg. The model is based on LLaMA architecture extended with multimodal capabilities for vision-language understanding and localization.

Frequently asked

What is FoundationVision/Groma?
Groma is a multimodal LLM that uses localized visual tokenization to enable region-level understanding and visual grounding capabilities.
Is Groma open source?
Yes — FoundationVision/Groma is open source, released under the Apache-2.0 license.
What language is Groma written in?
FoundationVision/Groma is primarily written in Python.
How popular is Groma?
FoundationVision/Groma has 585 stars on GitHub.
Where can I find Groma?
FoundationVision/Groma is on GitHub at https://github.com/FoundationVision/Groma.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.