A Python utility that converts office documents and media into structured Markdown built for LLM pipelines, not human eyeballs.
Data Tooling
big names on the moveFirecrawl turns search, scraping, and browser interaction into a single API so your agents can read the web without wrestling with proxies, rate limits, or JavaScript rendering.
It turns technical PDFs and EPUBs into on-demand Claude Code skills, so you can query a book's actual frameworks instead of hoping the model remembers them.
Built in a fit of rage after a $16 “open source” web-to-Markdown tool gated features behind API tokens.
MinerU turns PDFs, Office files, and images into structured Markdown and JSON so LLM agents don’t drown in layout noise.
It turns images and PDFs into structured JSON and Markdown so your RAG pipeline doesn't have to squint.
ScrapeGraphAI lets you extract structured data from websites and documents by describing what you want in plain English, leaving the LLM to wrestle with the markup.
Docling turns PDFs, Office files, images, and even audio into structured AI-ready formats, entirely on your own hardware.
Built to stop pipelines from sending text-based PDFs through expensive OCR services.
OpenBB normalizes proprietary and public financial data so engineers can feed the same sources to Python scripts, REST APIs, Excel, and AI agents without rebuilding integrations.
It translates scientific papers into bilingual PDFs while keeping formulas, charts, and annotations exactly where they belong.
A maintainer cataloged every Chinese NLP repo they touched into a single, obsessively categorized list so others wouldn’t have to hunt.
OpenDataLoader PDF exists to extract structured data from PDFs for AI pipelines while auto-tagging untagged documents for screen readers, all without proprietary dependencies.
Repomix exists because copy-pasting twenty files into a chat window is a terrible way to ask an LLM for help.
It exists to handle the tedious wiring—annotations, dataset formats, tracking—that sits between a trained model and a useful application.
Maxun is an open-source platform for developers who would rather record a browsing session than write another brittle web scraper.
Chroma is an open-source search backend that handles the messy embedding pipeline so AI applications can store and retrieve documents with a minimal API.
Because someone has to label the training data, and it might as well not be in a spreadsheet.
LangExtract exists because asking an LLM to pull names and dates out of a report is easy; proving exactly which sentence each came from is the hard part.
It gives retail traders a local GUI for stitching together data feeds, backtesters, and execution frameworks into an AI-themed quant workflow.


