← all repositories

double22a/speech_dataset

A curated list of Chinese and English speech recognition datasets with durations and download links.

461 stars Data Tooling
speech_dataset
Not currently ranked — collecting fresh signals.
star history

This repository aggregates and documents publicly available speech datasets for automatic speech recognition research and development. It catalogs datasets across multiple languages including Mandarin Chinese and English, with metadata such as duration in hours and source URLs. The listed datasets include well-known resources like LibriSpeech, Common Voice, Aishell, and WenetSpeech, serving as a reference index for speech ML practitioners.

Frequently asked

What is double22a/speech_dataset?
A curated list of Chinese and English speech recognition datasets with durations and download links.
Is speech_dataset open source?
Yes — double22a/speech_dataset is open source, released under the Apache-2.0 license.
How popular is speech_dataset?
double22a/speech_dataset has 461 stars on GitHub.
Where can I find speech_dataset?
double22a/speech_dataset is on GitHub at https://github.com/double22a/speech_dataset.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.