← all repositories

chaoswork/sft_datasets

A curated repository of open-source datasets for supervised fine-tuning of large language models.

sft_datasets
Not currently ranked — collecting fresh signals.
star history

This repository organizes and catalogs open-source SFT datasets used for fine-tuning large language models, specifically focusing on Chinese language data. It includes diverse datasets for instruction following, mathematical reasoning, dialogue generation, and multi-task NLP. Each entry documents the dataset size, language, task type, generation method, and download links to sources like Hugging Face.

Frequently asked

What is chaoswork/sft_datasets?
A curated repository of open-source datasets for supervised fine-tuning of large language models.
Is sft_datasets open source?
Yes — chaoswork/sft_datasets is an open-source project tracked on heatdrop.
How popular is sft_datasets?
chaoswork/sft_datasets has 583 stars on GitHub.
Where can I find sft_datasets?
chaoswork/sft_datasets is on GitHub at https://github.com/chaoswork/sft_datasets.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.