← all repositories
edyoda/data-science-complete-tutorial

Classical ML bootcamp in a box: sixteen lessons, eight case studies

Jupyter notebooks that serve as the open companion to an instructor-led data-science program, walking through the scikit-learn ecosystem from NumPy to ensemble methods.

1.8k stars Jupyter Notebook LearningML Frameworks
data-science-complete-tutorial
Not currently ranked — collecting fresh signals.
star history

What it does This repo is essentially a digital textbook. Sixteen sequential Jupyter notebooks walk through the standard scikit-learn stack—NumPy, Pandas, preprocessing, linear models, decision trees, SVMs, clustering, and more—followed by eight applied case studies ranging from cancer prediction to customer churn. It functions as the free companion codebook for EdYoda’s data-science program.

The interesting bit Rather than a software library, this is curriculum-as-repository. The maintainers treat GitHub as a learning-management system, using nothing but Markdown links and notebook files to deliver a full classical-ML syllabus. It is an unusually complete, linear progression through pre-deep-learning fundamentals.

Key highlights

  • Sixteen topic notebooks from NumPy basics through ensemble methods and anomaly detection
  • Eight end-to-end case studies including income prediction, employee exit forecasting, and face generation
  • Explicit coverage of often-skipped practical topics: imbalanced classes, feature selection, and composite Pipeline estimators
  • Tied to video lectures and an external program, but the notebooks themselves require no login
  • 1,800+ stars suggest it has found an audience beyond the original classroom

Caveats

  • The README is a bare table of contents; there is no narrative guidance on how to navigate the sequence or prerequisites
  • Some notebook links still point to the old zekelabs organization, hinting the repo may not be actively maintained
  • Case-study depth is unclear from the README alone; titles promise a lot, but you cannot tell which techniques each project actually uses without opening the files

Verdict A solid starting point for self-learners who want a structured, classical-ML syllabus without video fluff. Skip it if you are looking for a reusable library, modern deep-learning frameworks, or interactive exercises beyond static notebooks.

Frequently asked

What is edyoda/data-science-complete-tutorial?
Jupyter notebooks that serve as the open companion to an instructor-led data-science program, walking through the scikit-learn ecosystem from NumPy to ensemble methods.
Is data-science-complete-tutorial open source?
Yes — edyoda/data-science-complete-tutorial is an open-source project tracked on heatdrop.
What language is data-science-complete-tutorial written in?
edyoda/data-science-complete-tutorial is primarily written in Jupyter Notebook.
How popular is data-science-complete-tutorial?
edyoda/data-science-complete-tutorial has 1.8k stars on GitHub.
Where can I find data-science-complete-tutorial?
edyoda/data-science-complete-tutorial is on GitHub at https://github.com/edyoda/data-science-complete-tutorial.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.