← all repositories

stepfun-ai/Step-Audio-EditX

A 3B-parameter LLM-based audio editing model that controls emotion, speaking style, and paralinguistics via reinforcement learning.

951 stars Python Image · Video · Audio
Step-Audio-EditX
Not currently ranked — collecting fresh signals.
star history

Step-Audio-EditX is a large language model for audio editing and synthesis. It enables editing of emotion, speaking style, and paralinguistic features in audio while supporting zero-shot text-to-speech. The model is trained using reinforcement learning techniques including SFT, DPO, and GRPO. It supports cross-lingual capabilities including English, Japanese, and Korean, and can be deployed via vLLM for efficient inference.

Frequently asked

What is stepfun-ai/Step-Audio-EditX?
A 3B-parameter LLM-based audio editing model that controls emotion, speaking style, and paralinguistics via reinforcement learning.
Is Step-Audio-EditX open source?
Yes — stepfun-ai/Step-Audio-EditX is open source, released under the Apache-2.0 license.
What language is Step-Audio-EditX written in?
stepfun-ai/Step-Audio-EditX is primarily written in Python.
How popular is Step-Audio-EditX?
stepfun-ai/Step-Audio-EditX has 951 stars on GitHub.
Where can I find Step-Audio-EditX?
stepfun-ai/Step-Audio-EditX is on GitHub at https://github.com/stepfun-ai/Step-Audio-EditX.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.