X-LANCE/SLAM-LLM
A training framework and toolkit for building custom multimodal LLMs that process speech, language, audio, and music.

SLAM-LLM is a deep learning framework that enables researchers and developers to train custom multimodal large language models (MLLMs) for speech, language, audio, and music processing tasks. It provides training recipes, PEFT support for efficient fine-tuning, and high-performance inference checkpoints. The framework supports multi-task training for ASR and speech translation, and scales to datasets with hundreds of thousands of hours of speech data.
Frequently asked
- What is X-LANCE/SLAM-LLM?
- A training framework and toolkit for building custom multimodal LLMs that process speech, language, audio, and music.
- Is SLAM-LLM open source?
- Yes — X-LANCE/SLAM-LLM is open source, released under the MIT license.
- What language is SLAM-LLM written in?
- X-LANCE/SLAM-LLM is primarily written in Python.
- How popular is SLAM-LLM?
- X-LANCE/SLAM-LLM has 1k stars on GitHub.
- Where can I find SLAM-LLM?
- X-LANCE/SLAM-LLM is on GitHub at https://github.com/X-LANCE/SLAM-LLM.