Zyphra/Zonos
An open-weight text-to-speech model trained on 200k+ hours of multilingual speech with speaker cloning and emotion control capabilities.

Not currently ranked — collecting fresh signals.
star history
Zonos-v0.1 is an open-weight TTS model that generates natural speech from text prompts using speaker embeddings or reference audio clips. It leverages a transformer/hybrid backbone with eSpeak-based text normalization and DAC token prediction. The model supports speech cloning from short reference clips and fine-grained control over speaking rate, pitch, audio quality, and emotional expression, outputting natively at 44kHz.
Frequently asked
- What is Zyphra/Zonos?
- An open-weight text-to-speech model trained on 200k+ hours of multilingual speech with speaker cloning and emotion control capabilities.
- Is Zonos open source?
- Yes — Zyphra/Zonos is open source, released under the Apache-2.0 license.
- What language is Zonos written in?
- Zyphra/Zonos is primarily written in Python.
- How popular is Zonos?
- Zyphra/Zonos has 7.2k stars on GitHub.
- Where can I find Zonos?
- Zyphra/Zonos is on GitHub at https://github.com/Zyphra/Zonos.