← all repositories

lyuchenyang/Macaw-LLM

Multi-modal LLM combining vision, audio, and text processing for unified language modeling.

Macaw-LLM
Not currently ranked — collecting fresh signals.
star history

Macaw-LLM is a research project developing multi-modal language modeling capabilities by integrating images, videos, audio, and text into a unified system. The architecture leverages pre-trained components including CLIP for visual understanding, Whisper for audio processing, and LLaMA as the base language model. This enables the model to process and reason across multiple modalities within a language modeling framework.

Frequently asked

What is lyuchenyang/Macaw-LLM?
Multi-modal LLM combining vision, audio, and text processing for unified language modeling.
Is Macaw-LLM open source?
Yes — lyuchenyang/Macaw-LLM is open source, released under the Apache-2.0 license.
What language is Macaw-LLM written in?
lyuchenyang/Macaw-LLM is primarily written in Python.
How popular is Macaw-LLM?
lyuchenyang/Macaw-LLM has 1.6k stars on GitHub.
Where can I find Macaw-LLM?
lyuchenyang/Macaw-LLM is on GitHub at https://github.com/lyuchenyang/Macaw-LLM.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.