OpenMOSS/MOSS-TTSD
A multi-speaker text-to-speech model for expressive spoken dialogue generation with zero-shot voice cloning from short audio references.

Not currently ranked — collecting fresh signals.
star history
MOSS-TTSD is a spoken dialogue generation model designed for expressive multi-speaker synthesis. It features long-context modeling, flexible speaker control, and multilingual support, while enabling zero-shot voice cloning from short audio references. The model is built with PyTorch and supports fine-tuning for real-world long-form content creation including podcasts, audiobooks, and entertainment scenarios.
Frequently asked
- What is OpenMOSS/MOSS-TTSD?
- A multi-speaker text-to-speech model for expressive spoken dialogue generation with zero-shot voice cloning from short audio references.
- Is MOSS-TTSD open source?
- Yes — OpenMOSS/MOSS-TTSD is open source, released under the Apache-2.0 license.
- What language is MOSS-TTSD written in?
- OpenMOSS/MOSS-TTSD is primarily written in Python.
- How popular is MOSS-TTSD?
- OpenMOSS/MOSS-TTSD has 1.4k stars on GitHub.
- Where can I find MOSS-TTSD?
- OpenMOSS/MOSS-TTSD is on GitHub at https://github.com/OpenMOSS/MOSS-TTSD.