It wants to parse entire documents in one shot without the model getting stuck in repetitive loops.
Computer Vision
underdogs breaking outIt reconstructs 3D scenes from streaming video in real time without per-scene optimization, using a feed-forward transformer that remembers trajectory and corrects drift as it goes.
It removes the visible Gemini sparkle, invisible SynthID fingerprints, and C2PA metadata that AI image generators embed in every output.
LaTeXSnipper bundles screenshot OCR, handwriting recognition, symbolic computation, and Office plugins into a single offline desktop app so you can actually use the math you capture.
TurboOCR exists because waiting for a vision-language model to read a receipt is a waste of GPU time.
An open 10B-parameter VLA model that predicts driving trajectories while spelling out the causal reasoning behind every turn and lane change.
HunyuanOCR-1.5 speeds up lightweight OCR vision-language models by drafting tokens with a block-diffusion model and closing capability gaps with an agent-driven data pipeline.
AlpaSim exists so researchers can validate end-to-end self-driving policies in a closed-loop Python sandbox where renderers, physics, and traffic are swappable microservices.
RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.
Built to let you swap out the plate detector and OCR engine without trashing the rest of the pipeline.
It converts images and PDFs into structured HTML, Markdown, or JSON while reconstructing tables, forms, and handwriting that most OCR tools reduce to plain text soup.
It gives Python developers on Windows and Linux a fully offline shortcut for detecting faces, extracting landmarks, and scoring similarity.
Eagle is less a single model than NVIDIA's internal R&D pipeline for multimodal AI, now open-sourced with three generations of VLMs and a grounding specialist.
screenpipe continuously records your screen and audio locally so AI can search, summarize, and act on everything you’ve done without sending data to the cloud.
It locally inpaints over hard-coded subtitles and text watermarks in videos and images so you never have to upload frames to a cloud API.
This tool automates image and video annotation by plugging dozens of SOTA models into a single PyQt6 GUI.
VGGT replaces the traditional multi-stage 3D reconstruction pipeline with a single feed-forward model that predicts cameras, depth, and geometry from one or many images in seconds.
An ONNX-exported, multi-engine OCR toolkit that runs offline on basically anything.
Koharu automates the full manga-translation pipeline—detection, OCR, inpainting, and text rendering—entirely on your local machine.
A single C++17 API wraps detection, segmentation, pose, OBB, and classification across YOLOv5 through YOLO26, no Python runtime required.




