YuanGongND/ltu
An audio and speech large language model that bridges audio/speech perception with natural language understanding capabilities.

LTU and LTU-AS are audio and speech large language models that process audio input to enable open-ended question answering alongside strong performance on closed-ended audio tasks. The repository provides PyTorch implementations with pretrained checkpoints, datasets (OpenAQA and OpenASQA), training reproduction code, and fine-tuning capabilities. Interactive HuggingFace Space demos allow users to interact with the models without local GPU resources.
Frequently asked
- What is YuanGongND/ltu?
- An audio and speech large language model that bridges audio/speech perception with natural language understanding capabilities.
- Is ltu open source?
- Yes — YuanGongND/ltu is an open-source project tracked on heatdrop.
- What language is ltu written in?
- YuanGongND/ltu is primarily written in Python.
- How popular is ltu?
- YuanGongND/ltu has 477 stars on GitHub.
- Where can I find ltu?
- YuanGongND/ltu is on GitHub at https://github.com/YuanGongND/ltu.