Token Monitor reads local logs from two dozen AI coding tools to surface live token burn, costs, and limits in one place, synced across all your machines.
Data Tooling
underdogs · picking up speedBecause your coding assistant shouldn't need a securities API cheat sheet to analyze A-shares.
Turns URLs into clean markdown and JSON so AI agents don't have to parse HTML soup.
It turns technical PDFs and EPUBs into on-demand Claude Code skills, so you can query a book's actual frameworks instead of hoping the model remembers them.
It removes the visible Gemini sparkle, invisible SynthID fingerprints, and C2PA metadata that AI image generators embed in every output.
Uses multimodal LLMs to transcribe PDFs into Markdown, preserving complex layouts that traditional extractors mangle.
It exists to automate the tedious pipeline of turning noisy PDFs and plain text into structured training data for domain-specific LLMs.
OpenLake wants storage to bypass the host entirely and land straight in GPU memory.
This unofficial rebuild of a Google Research project chains specialized agents to turn rough text and data into publication-ready academic figures.
Unstract turns document extraction into a prompt-and-deploy workflow instead of a regex archaeology dig.
A pipeline that turns messy PDFs and slides into structured, navigable memory for AI agents instead of flat text shards.
ktx is a local context layer that ingests your data stack and business knowledge so Claude, Codex, and other agents query warehouses with approved metrics instead of inventing SQL.
It scrapes WeChat public articles and serves them as RSS, Markdown, and JSON because Tencent won’t.
It turns Git repositories into flat, token-counted text digests so you can stop manually concatenating files for LLM prompts.
It captures the messy reality of long-horizon agent tasks and turns execution traces into reusable, shareable learning signals.
Maxun is an open-source platform for developers who would rather record a browsing session than write another brittle web scraper.
Open Wearables wants to unify fitness tracker data behind one self-hosted API so developers can stop writing bespoke OAuth flows for every wearable brand.
olmOCR exists because LLMs cannot train on PDFs until someone strips the formatting chaos and restores natural reading order.
This tool automates image and video annotation by plugging dozens of SOTA models into a single PyQt6 GUI.
Someone finally collected all the ML-for-security papers, datasets, and books in one place so you don't have to hunt through conference proceedings at 2 AM.




