HazyResearch/data-centric-ai
A curated collection of resources, papers, and guides about data-centric AI methodologies from Stanford's HazyResearch group.

This repository consolidates resources on data-centric AI, an approach that prioritizes improving data quality over model architecture. It collects papers, blog posts, and research progress in techniques such as data labeling, curation, and validation that help ML practitioners achieve better real-world results. The project originated from Stanford’s HazyResearch group and includes contributions from the broader data-centric AI community.
Frequently asked
- What is HazyResearch/data-centric-ai?
- A curated collection of resources, papers, and guides about data-centric AI methodologies from Stanford's HazyResearch group.
- Is data-centric-ai open source?
- Yes — HazyResearch/data-centric-ai is open source, released under the Apache-2.0 license.
- What language is data-centric-ai written in?
- HazyResearch/data-centric-ai is primarily written in TeX.
- How popular is data-centric-ai?
- HazyResearch/data-centric-ai has 1.1k stars on GitHub.
- Where can I find data-centric-ai?
- HazyResearch/data-centric-ai is on GitHub at https://github.com/HazyResearch/data-centric-ai.