zai-org/GLM-TTS
A text-to-speech synthesis system using large language models that supports zero-shot voice cloning and emotion control via multi-reward reinforcement learning.

Not currently ranked — collecting fresh signals.
star history
GLM-TTS is a high-quality TTS system based on large language models with a two-stage architecture: an LLM generates speech token sequences and a Flow model converts them to audio waveforms. It introduces multi-reward reinforcement learning for improved emotional expression and natural prosody control, supporting zero-shot voice cloning with 3-10 seconds of prompt audio and streaming inference.
Frequently asked
- What is zai-org/GLM-TTS?
- A text-to-speech synthesis system using large language models that supports zero-shot voice cloning and emotion control via multi-reward reinforcement learning.
- Is GLM-TTS open source?
- Yes — zai-org/GLM-TTS is open source, released under the Apache-2.0 license.
- What language is GLM-TTS written in?
- zai-org/GLM-TTS is primarily written in Python.
- How popular is GLM-TTS?
- zai-org/GLM-TTS has 1k stars on GitHub.
- Where can I find GLM-TTS?
- zai-org/GLM-TTS is on GitHub at https://github.com/zai-org/GLM-TTS.