wasiahmad/Awesome-LLM-Synthetic-Data
A curated reading list of papers, tools, and blogs on using LLMs to generate synthetic data for training and improving language models.

Not currently ranked — collecting fresh signals.
star history
This repository compiles research and resources on synthetic data generation using large language models. It covers methods and applications across mathematical reasoning, code generation, alignment, reward modeling, long-context understanding, multi-modal tasks, and agent systems. Organized as an awesome-style list, it serves as a reference for researchers and practitioners working on LLM training data creation.
Frequently asked
- What is wasiahmad/Awesome-LLM-Synthetic-Data?
- A curated reading list of papers, tools, and blogs on using LLMs to generate synthetic data for training and improving language models.
- Is Awesome-LLM-Synthetic-Data open source?
- Yes — wasiahmad/Awesome-LLM-Synthetic-Data is open source, released under the MIT license.
- How popular is Awesome-LLM-Synthetic-Data?
- wasiahmad/Awesome-LLM-Synthetic-Data has 1.5k stars on GitHub.
- Where can I find Awesome-LLM-Synthetic-Data?
- wasiahmad/Awesome-LLM-Synthetic-Data is on GitHub at https://github.com/wasiahmad/Awesome-LLM-Synthetic-Data.