Built to stop pipelines from sending text-based PDFs through expensive OCR services.
Data Tooling
underdogs · picking up speedToken Monitor reads local logs from two dozen AI coding tools to surface live token burn, costs, and limits in one place, synced across all your machines.
omm uses your AI coding tool to scan codebases and emit nested Mermaid diagrams, then renders them in a local viewer.
Datus is an open-source agent that keeps LLMs from hallucinating SQL by building a living, learning knowledge base around your data stack.
A thin Python wrapper around PyMuPDF that turns documents into structured Markdown, JSON, or plain text—layout-aware, with selective OCR that skips clean pages.
Because your coding assistant shouldn't need a securities API cheat sheet to analyze A-shares.
Most AI image prompts are one-off text blobs; this repo distills 96 visual styles into structured JSON templates so you can swap variables without losing style direction.
RoboTwin 2.0 generates synthetic training data and benchmarks for bimanual manipulation, because two arms are harder than one and real-world demos are expensive.
It wraps IBM's Docling engine in a FastAPI service backed by Celery and Redis, giving you sync, async, and batch endpoints for heavy-duty document-to-Markdown conversion.
Built to give AI agents a faster, more accurate way to turn the web into markdown, offering drop-in Firecrawl compatibility and a self-hostable single binary.
It exists because hand-coding knowledge graphs from PDFs and YouTube transcripts is nobody's idea of a good time.
A 5.7k-star duplication detector rebuilt itself for the agentic era: token-efficient reporters, MCP server, and skills your AI assistant can actually use.
A Dart SDK that lets Flutter apps upload video and search inside it by meaning, speech, or imagery—without imposing any UI.
Kreuzberg wraps a Rust document engine in native bindings for more than a dozen languages so you can extract text and metadata from 90+ formats without switching stacks.
A Python tool that turns noisy HTML into clean, structured text for NLP pipelines and research corpora.
Modular agent skills that hand LLMs the boring office work—Excel analysis, slide decks, research reports, and infographics—instead of letting them hallucinate about it.
Scans running WeChat process memory to extract SQLCipher keys, then exposes your chat database as a JSON-first CLI designed for AI agent consumption.
A library that turns 'download this 50GB dataset' into a one-liner with streaming, caching, and zero-copy memory mapping.
An MCP server that searches 20+ academic sources and actually tells you when it can't download something instead of hallucinating a PDF.
OpenMetadata exists because plugging AI into a database gives it raw tables but no idea whether `cust_id` means customer, account, or buyer.



