← all repositories

kyutai-labs/delayed-streams-modeling

Kyutai Labs' speech-to-text and text-to-speech models using Delayed Streams Modeling for real-time streaming audio processing.

delayed-streams-modeling
Not currently ranked — collecting fresh signals.
star history

This repository provides Kyutai’s STT and TTS models based on the Delayed Streams Modeling (DSM) framework, a technique for streaming speech-to-text and text-to-speech tasks. The STT models include a 1B parameter English/French model with 0.5 second delay and a 2.6B English-only model with 2.5 second delay. Both models support real-time streaming inference, efficient batching (400 streams per H100), and word-level timestamps. The framework is documented in a pre-print at arxiv.org/abs/2509.08753.

Frequently asked

What is kyutai-labs/delayed-streams-modeling?
Kyutai Labs' speech-to-text and text-to-speech models using Delayed Streams Modeling for real-time streaming audio processing.
Is delayed-streams-modeling open source?
Yes — kyutai-labs/delayed-streams-modeling is open source, released under the Apache-2.0 license.
What language is delayed-streams-modeling written in?
kyutai-labs/delayed-streams-modeling is primarily written in Python.
How popular is delayed-streams-modeling?
kyutai-labs/delayed-streams-modeling has 3k stars on GitHub.
Where can I find delayed-streams-modeling?
kyutai-labs/delayed-streams-modeling is on GitHub at https://github.com/kyutai-labs/delayed-streams-modeling.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.