chaoswork/sft_datasets
A curated repository of open-source datasets for supervised fine-tuning of large language models.

Not currently ranked — collecting fresh signals.
star history
This repository organizes and catalogs open-source SFT datasets used for fine-tuning large language models, specifically focusing on Chinese language data. It includes diverse datasets for instruction following, mathematical reasoning, dialogue generation, and multi-task NLP. Each entry documents the dataset size, language, task type, generation method, and download links to sources like Hugging Face.
Frequently asked
- What is chaoswork/sft_datasets?
- A curated repository of open-source datasets for supervised fine-tuning of large language models.
- Is sft_datasets open source?
- Yes — chaoswork/sft_datasets is an open-source project tracked on heatdrop.
- How popular is sft_datasets?
- chaoswork/sft_datasets has 583 stars on GitHub.
- Where can I find sft_datasets?
- chaoswork/sft_datasets is on GitHub at https://github.com/chaoswork/sft_datasets.