double22a/speech_dataset
A curated list of Chinese and English speech recognition datasets with durations and download links.
★461 stars Data Tooling

Not currently ranked — collecting fresh signals.
star history
This repository aggregates and documents publicly available speech datasets for automatic speech recognition research and development. It catalogs datasets across multiple languages including Mandarin Chinese and English, with metadata such as duration in hours and source URLs. The listed datasets include well-known resources like LibriSpeech, Common Voice, Aishell, and WenetSpeech, serving as a reference index for speech ML practitioners.
Frequently asked
- What is double22a/speech_dataset?
- A curated list of Chinese and English speech recognition datasets with durations and download links.
- Is speech_dataset open source?
- Yes — double22a/speech_dataset is open source, released under the Apache-2.0 license.
- How popular is speech_dataset?
- double22a/speech_dataset has 461 stars on GitHub.
- Where can I find speech_dataset?
- double22a/speech_dataset is on GitHub at https://github.com/double22a/speech_dataset.