← all repositories

kyegomez/MultiModalMamba

Multi-modal deep learning model combining Vision Transformer and Mamba SSM architectures for concurrent text and image processing.

MultiModalMamba
Not currently ranked — collecting fresh signals.
star history

MultiModalMamba implements a novel architecture fusing Vision Transformer (ViT) with Mamba state space models to create a high-performance multi-modal model. The architecture processes both text sequences and images concurrently, using transformer attention mechanisms alongside efficient Mamba layers for feature extraction and fusion. Built on Zeta, a minimalist PyTorch-based AI framework, the model provides a MultiModalMambaBlock component and a full trainable model for multi-modal tasks.

Frequently asked

What is kyegomez/MultiModalMamba?
Multi-modal deep learning model combining Vision Transformer and Mamba SSM architectures for concurrent text and image processing.
Is MultiModalMamba open source?
Yes — kyegomez/MultiModalMamba is open source, released under the MIT license.
What language is MultiModalMamba written in?
kyegomez/MultiModalMamba is primarily written in Python.
How popular is MultiModalMamba?
kyegomez/MultiModalMamba has 473 stars on GitHub.
Where can I find MultiModalMamba?
kyegomez/MultiModalMamba is on GitHub at https://github.com/kyegomez/MultiModalMamba.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.