ictnlp/StreamSpeech
StreamSpeech is a multi-task neural model that performs simultaneous speech-to-text and speech-to-speech translation alongside speech synthesis.

StreamSpeech is an end-to-end speech processing model that unifies automatic speech recognition, machine translation, and speech synthesis in a single architecture. It supports both offline and streaming (simultaneous) modes for speech-to-text and speech-to-speech translation tasks. The model is trained using multi-task learning and achieves state-of-the-art results on both offline and simultaneous translation benchmarks.
Frequently asked
- What is ictnlp/StreamSpeech?
- StreamSpeech is a multi-task neural model that performs simultaneous speech-to-text and speech-to-speech translation alongside speech synthesis.
- Is StreamSpeech open source?
- Yes — ictnlp/StreamSpeech is open source, released under the MIT license.
- What language is StreamSpeech written in?
- ictnlp/StreamSpeech is primarily written in Python.
- How popular is StreamSpeech?
- ictnlp/StreamSpeech has 1.3k stars on GitHub.
- Where can I find StreamSpeech?
- ictnlp/StreamSpeech is on GitHub at https://github.com/ictnlp/StreamSpeech.