Data Tooling

Data Tooling

big names · picking up speed
01
microsoft/markitdown
+635 ★/dayaccelerating

A Python utility that converts office documents and media into structured Markdown built for LLM pipelines, not human eyeballs.

182.5k Python Data Tooling · explained Feature
02
opendatalab/MinerU
+78 ★/dayaccelerating

MinerU turns PDFs, Office files, and images into structured Markdown and JSON so LLM agents don’t drown in layout noise.

79.6k Python Data Tooling · explained
03

OpenDataLoader PDF exists to extract structured data from PDFs for AI pipelines while auto-tagging untagged documents for screen readers, all without proprietary dependencies.

29.1k Java Data Tooling · explained
04
fighting41love/funNLP
+23 ★/dayaccelerating

A maintainer cataloged every Chinese NLP repo they touched into a single, obsessively categorized list so others wouldn’t have to hunt.

83k Python Learning · explained
05
HumanSignal/label-studio
+7.3 ★/dayaccelerating

Because someone has to label the training data, and it might as well not be in a spreadsheet.

28.2k TypeScript Data Tooling · explained
06
google/langextract
+5.9 ★/dayaccelerating

LangExtract exists because asking an LLM to pull names and dates out of a report is easy; proving exactly which sentence each came from is the hard part.

38.6k Python Data Tooling · explained
07
getmaxun/maxun
+8.1 ★/dayaccelerating

Maxun is an open-source platform for developers who would rather record a browsing session than write another brittle web scraper.

17.4k TypeScript Data Tooling · explained
08
chroma-core/chroma
+7.9 ★/daysteady

Chroma is an open-source search backend that handles the messy embedding pipeline so AI applications can store and retrieve documents with a minimal API.

29.3k Rust RAG · Search · explained
09
UFund-Me/Qbot
+5.3 ★/daysteady

It gives retail traders a local GUI for stitching together data feeds, backtesters, and execution frameworks into an AI-themed quant workflow.

18.5k Jupyter Notebook Domain Apps · explained
11
yamadashy/repomix
+15 ★/daycooling

Repomix exists because copy-pasting twenty files into a chat window is a terrible way to ask an LLM for help.

28.3k TypeScript Data Tooling · explained
12
roboflow/supervision
+13 ★/daycooling

It exists to handle the tedious wiring—annotations, dataset formats, tracking—that sits between a trained model and a useful application.

50k Python Computer Vision · explained Feature
13
toon-format/toon
+5.0 ★/daycooling

It re-encodes JSON into a token-cheaper, schema-explicit format so you can fit more context into LLM prompts without losing structure.

25.4k TypeScript LLMOps · Eval · explained
14
academic/awesome-datascience
+4.6 ★/daycooling

A curated awesome-list that tries to answer "What is Data Science, and what should I study?" by cataloging courses, tools, libraries, and communities in a single sprawling index.

30k Learning · explained
16
OpenBB-finance/OpenBB
+32 ★/daycooling

OpenBB normalizes proprietary and public financial data so engineers can feed the same sources to Python scripts, REST APIs, Excel, and AI agents without rebuilding integrations.

72.9k Python Domain Apps · explained
17

It translates scientific papers into bilingual PDFs while keeping formulas, charts, and annotations exactly where they belong.

36.8k Python Other AI · explained
18
ScrapeGraphAI/Scrapegraph-ai
+46 ★/daycooling

ScrapeGraphAI lets you extract structured data from websites and documents by describing what you want in plain English, leaving the LLM to wrestle with the markup.

30.8k Python RAG · Search · explained
19
firecrawl/firecrawl
+383 ★/daycooling

Firecrawl turns search, scraping, and browser interaction into a single API so your agents can read the web without wrestling with proxies, rate limits, or JavaScript rendering.

178.9k TypeScript Data Tooling · explained Feature
20
virgiliojr94/book-to-skill
+213 ★/daycooling

It turns technical PDFs and EPUBs into on-demand Claude Code skills, so you can query a book's actual frameworks instead of hoping the model remembers them.

29.8k Python Coding Assistants · explained Feature
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.