← all repositories

zhenye234/LLaSA_training

A speech synthesis model built on LLaMA architecture that generates audio from text using scaled train-time and inference-time compute.

LLaSA_training
Not currently ranked — collecting fresh signals.
star history

LLaSA is an LLaMA-based neural text-to-speech system designed to generate natural speech from textual input. The system leverages scaled compute during both training and inference phases to improve output quality. It uses the XCodec2 codec for audio encoding and incorporates a Llama text tokenizer (e.g., Llama-3.2-1B-Instruct) for text encoding. Training supports distributed execution via torchrun or SLURM, and the project provides 160k hours of open-source tokenized speech data for training.

Frequently asked

What is zhenye234/LLaSA_training?
A speech synthesis model built on LLaMA architecture that generates audio from text using scaled train-time and inference-time compute.
Is LLaSA_training open source?
Yes — zhenye234/LLaSA_training is an open-source project tracked on heatdrop.
What language is LLaSA_training written in?
zhenye234/LLaSA_training is primarily written in Python.
How popular is LLaSA_training?
zhenye234/LLaSA_training has 660 stars on GitHub.
Where can I find LLaSA_training?
zhenye234/LLaSA_training is on GitHub at https://github.com/zhenye234/LLaSA_training.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.