EvolvingLMMs-Lab/Otter
Multi-modal LLM combining vision and language capabilities with instruction-following and in-context learning.

Not currently ranked — collecting fresh signals.
star history
Otter is an open-source multi-modal foundation model based on OpenFlamingo (itself based on DeepMind’s Flamingo). It is trained on the MIMIC-IT dataset and designed for instruction-following and in-context learning across vision-language tasks. The project provides model checkpoints on HuggingFace and supports both image and video understanding capabilities through specialized variants.
Frequently asked
- What is EvolvingLMMs-Lab/Otter?
- Multi-modal LLM combining vision and language capabilities with instruction-following and in-context learning.
- Is Otter open source?
- Yes — EvolvingLMMs-Lab/Otter is open source, released under the MIT license.
- What language is Otter written in?
- EvolvingLMMs-Lab/Otter is primarily written in Python.
- How popular is Otter?
- EvolvingLMMs-Lab/Otter has 3.4k stars on GitHub.
- Where can I find Otter?
- EvolvingLMMs-Lab/Otter is on GitHub at https://github.com/EvolvingLMMs-Lab/Otter.