It wants to parse entire documents in one shot without the model getting stuck in repetitive loops.
Computer Vision
big names · picking up speedIt turns face swaps and lip-syncs into queued, retryable batch jobs instead of one-off scripts.
MAA automates the daily chores of Arknights by treating the game screen as a computer vision problem, using OpenCV and OCR to handle farming, recruitment, and base shifts without human tapping.
It bundles detection, recognition, alignment, and reconstruction into a single research-grade toolbox.
It exists so you can extract text from screenshots, PDFs, and barcodes without a network connection or a cloud bill.
It turns images of text into searchable documents across more than 100 languages, offering both a command-line tool and a C++ library for builders.
OpenCV is an open-source C++ computer vision library whose own README acts as a portal rather than a product page.
Real-ESRGAN turns the ESRGAN research model into a practical tool for upscaling and restoring real-world images and videos using only synthetic training data.
It exists to let developers run customized vision, text, and audio machine learning across mobile, web, and edge hardware without cloud round-trips.
It turns images and PDFs into structured JSON and Markdown so your RAG pipeline doesn't have to squint.
A single library that collects, trains, and exports nearly every image backbone worth using—so you don't have to reimplement them yourself.
An open-source toolkit that turns casual selfies into compliant ID photos using offline AI matting and face detection.
screenpipe continuously records your screen and audio locally so AI can search, summarize, and act on everything you’ve done without sending data to the cloud.
It exists to handle the tedious wiring—annotations, dataset formats, tracking—that sits between a trained model and a useful application.
OCRmyPDF exists because most free OCR tools botch text placement, bloat file sizes, or mangle image resolution when trying to make scanned documents searchable.
Upscayl gives desktop users a free way to enlarge and enhance low-resolution images using local Real-ESRGAN models and a Vulkan-compatible GPU.
Ultralytics wants to stop you from stitching together separate repos for every computer vision task by bundling detection, segmentation, tracking, and pose estimation into one YOLO-backed package.
Google Research releases all its code and datasets in one place, which has grown so large that the README treats a full clone as a hazard.
Frigate performs real-time, local object detection on IP camera streams using OpenCV and TensorFlow, designed to integrate tightly with Home Assistant.
It rounds up hundreds of AI project links so you don't have to hunt them down yourself.



