← all repositories

wasiahmad/Awesome-LLM-Synthetic-Data

A curated reading list of papers, tools, and blogs on using LLMs to generate synthetic data for training and improving language models.

Awesome-LLM-Synthetic-Data
Not currently ranked — collecting fresh signals.
star history

This repository compiles research and resources on synthetic data generation using large language models. It covers methods and applications across mathematical reasoning, code generation, alignment, reward modeling, long-context understanding, multi-modal tasks, and agent systems. Organized as an awesome-style list, it serves as a reference for researchers and practitioners working on LLM training data creation.

Frequently asked

What is wasiahmad/Awesome-LLM-Synthetic-Data?
A curated reading list of papers, tools, and blogs on using LLMs to generate synthetic data for training and improving language models.
Is Awesome-LLM-Synthetic-Data open source?
Yes — wasiahmad/Awesome-LLM-Synthetic-Data is open source, released under the MIT license.
How popular is Awesome-LLM-Synthetic-Data?
wasiahmad/Awesome-LLM-Synthetic-Data has 1.5k stars on GitHub.
Where can I find Awesome-LLM-Synthetic-Data?
wasiahmad/Awesome-LLM-Synthetic-Data is on GitHub at https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.