It wants to parse entire documents in one shot without the model getting stuck in repetitive loops.
Computer Vision
underdogs · picking up speedIt removes the visible Gemini sparkle, invisible SynthID fingerprints, and C2PA metadata that AI image generators embed in every output.
VR eye cameras output flat images, and this project recovers the metric 3D geometry researchers actually need.
A testing ground where hard computer vision problems—ball tracking, jersey OCR, player re-ID—get solved with reusable tools.
It converts images and PDFs into structured HTML, Markdown, or JSON while reconstructing tables, forms, and handwriting that most OCR tools reduce to plain text soup.
A single C++17 API wraps detection, segmentation, pose, OBB, and classification across YOLOv5 through YOLO26, no Python runtime required.
Agent S is an open-source framework that lets autonomous AI agents operate a computer through its GUI, learning from experience to complete complex tasks.
Built to let you swap out the plate detector and OCR engine without trashing the rest of the pipeline.
This tool automates image and video annotation by plugging dozens of SOTA models into a single PyQt6 GUI.
Koharu automates the full manga-translation pipeline—detection, OCR, inpainting, and text rendering—entirely on your local machine.
RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.
MAA automates the daily chores of Arknights by treating the game screen as a computer vision problem, using OpenCV and OCR to handle farming, recruitment, and base shifts without human tapping.
VGGT replaces the traditional multi-stage 3D reconstruction pipeline with a single feed-forward model that predicts cameras, depth, and geometry from one or many images in seconds.
LichtFeld Studio wraps the entire 3D Gaussian Splatting pipeline—training, editing, exporting, automating—into a single C++ desktop app instead of a chain of Python scripts.
A local-only tool that turns burned-in video subtitles into clean SRT files using deep learning, no cloud OCR required.
TurboOCR exists because waiting for a vision-language model to read a receipt is a waste of GPU time.
An open 10B-parameter VLA model that predicts driving trajectories while spelling out the causal reasoning behind every turn and lane change.
HunyuanOCR-1.5 speeds up lightweight OCR vision-language models by drafting tokens with a block-diffusion model and closing capability gaps with an agent-driven data pipeline.
AlpaSim exists so researchers can validate end-to-end self-driving policies in a closed-loop Python sandbox where renderers, physics, and traffic are swappable microservices.
It gives Python developers on Windows and Linux a fully offline shortcut for detecting faces, extracting landmarks, and scoring similarity.




