← all repositories

paulpierre/markdown-crawler

A multithreaded web crawler that converts websites into markdown files for use in LLM RAG pipelines and knowledge bases.

455 stars Python Data ToolingRAG · Search
markdown-crawler
Not currently ranked — collecting fresh signals.
star history

This tool recursively crawls websites and generates markdown files for each page, preserving document structure like tables and images. It uses BeautifulSoup for HTML parsing and supports multithreading for faster crawling with resumable sessions. The output is designed to be easily chunked and processed for retrieval augmented generation systems, LLM fine-tuning datasets, and agent knowledge bases.

Frequently asked

What is paulpierre/markdown-crawler?
A multithreaded web crawler that converts websites into markdown files for use in LLM RAG pipelines and knowledge bases.
Is markdown-crawler open source?
Yes — paulpierre/markdown-crawler is open source, released under the MIT license.
What language is markdown-crawler written in?
paulpierre/markdown-crawler is primarily written in Python.
How popular is markdown-crawler?
paulpierre/markdown-crawler has 455 stars on GitHub.
Where can I find markdown-crawler?
paulpierre/markdown-crawler is on GitHub at https://github.com/paulpierre/markdown-crawler.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.