← all repositories
OpenDataBox/awesome-data-llm

The IaaS framework for high-quality LLM training data

This repository maps the papers and datasets behind the full data lifecycle for large language models, from web crawl to training batch.

awesome-data-llm
Not currently ranked — collecting fresh signals.
star history

What it does This repository serves as the official companion to a trio of survey papers on the intersection of large language models and data management. It catalogs relevant research, public datasets, and frameworks across the entire data lifecycle—from acquisition and deduplication to storage, serving, and even using LLMs as data analysts. The core organizing principle is the IaaS rubric: Inclusiveness, Abundance, Articulation, and Sanitization.

The interesting bit Most “awesome” lists are flat; this one is structured like a textbook, with sections for pre-training, SFT, RL, RAG, and agent data, plus a dedicated taxonomy for LLM-enhanced data preparation. The authors also propose the IaaS concept—phonetically echoing Infrastructure-as-a-Service—as a memorable four-pillar standard for dataset quality, which is a rare attempt to brand data curation principles.

Key highlights

  • Curated papers and datasets for every stage of the LLM lifecycle: pre-training, continual pre-training, SFT, RL, RAG, evaluation, and agents.
  • Introduces the IaaS framework (Inclusiveness, Abundance, Articulation, Sanitization) to define high-quality dataset characteristics.
  • Covers three distinct surveys: the main LLM × DATA survey, LLM/Agent-as-Data-Analyst, and LLM-enhanced application-ready data preparation.
  • Includes practical dataset pointers such as CommonCrawl, RedPajama, SlimPajama-627B-DC, Falcon-RefinedWeb, and domain-specific corpora like PubMed and Arxiv.
  • Organizes data management tasks into acquisition, deduplication, filtering, domain selection, mixing, distillation, synthesis, storage formats, and provenance.

Caveats

  • The README is primarily a table of contents and citation list; the deep synthesis lives in the PDF papers, not the repo itself.
  • Some overview images are hotlinked from other repositories rather than hosted locally.

Verdict Worth bookmarking if you are building training pipelines or studying data-centric AI. Skip it if you are looking for runnable code or tools, because this is a reading list and paper index, not a framework.

Frequently asked

What is OpenDataBox/awesome-data-llm?
This repository maps the papers and datasets behind the full data lifecycle for large language models, from web crawl to training batch.
Is awesome-data-llm open source?
Yes — OpenDataBox/awesome-data-llm is an open-source project tracked on heatdrop.
How popular is awesome-data-llm?
OpenDataBox/awesome-data-llm has 803 stars on GitHub.
Where can I find awesome-data-llm?
OpenDataBox/awesome-data-llm is on GitHub at https://github.com/OpenDataBox/awesome-data-llm.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.