lyuchenyang/Macaw-LLM
Multi-modal LLM combining vision, audio, and text processing for unified language modeling.

Not currently ranked — collecting fresh signals.
star history
Macaw-LLM is a research project developing multi-modal language modeling capabilities by integrating images, videos, audio, and text into a unified system. The architecture leverages pre-trained components including CLIP for visual understanding, Whisper for audio processing, and LLaMA as the base language model. This enables the model to process and reason across multiple modalities within a language modeling framework.
Frequently asked
- What is lyuchenyang/Macaw-LLM?
- Multi-modal LLM combining vision, audio, and text processing for unified language modeling.
- Is Macaw-LLM open source?
- Yes — lyuchenyang/Macaw-LLM is open source, released under the Apache-2.0 license.
- What language is Macaw-LLM written in?
- lyuchenyang/Macaw-LLM is primarily written in Python.
- How popular is Macaw-LLM?
- lyuchenyang/Macaw-LLM has 1.6k stars on GitHub.
- Where can I find Macaw-LLM?
- lyuchenyang/Macaw-LLM is on GitHub at https://github.com/lyuchenyang/Macaw-LLM.