It wants to parse entire documents in one shot without the model getting stuck in repetitive loops.
Computer Vision
underdogs breaking outIt reconstructs 3D scenes from streaming video in real time without per-scene optimization, using a feed-forward transformer that remembers trajectory and corrects drift as it goes.
It removes the visible Gemini sparkle, invisible SynthID fingerprints, and C2PA metadata that AI image generators embed in every output.
LaTeXSnipper bundles screenshot OCR, handwriting recognition, symbolic computation, and Office plugins into a single offline desktop app so you can actually use the math you capture.
TurboOCR exists because waiting for a vision-language model to read a receipt is a waste of GPU time.
An open 10B-parameter VLA model that predicts driving trajectories while spelling out the causal reasoning behind every turn and lane change.
HunyuanOCR-1.5 speeds up lightweight OCR vision-language models by drafting tokens with a block-diffusion model and closing capability gaps with an agent-driven data pipeline.
AlpaSim exists so researchers can validate end-to-end self-driving policies in a closed-loop Python sandbox where renderers, physics, and traffic are swappable microservices.
A testing ground where hard computer vision problems—ball tracking, jersey OCR, player re-ID—get solved with reusable tools.
It converts images and PDFs into structured HTML, Markdown, or JSON while reconstructing tables, forms, and handwriting that most OCR tools reduce to plain text soup.
RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.
Built to let you swap out the plate detector and OCR engine without trashing the rest of the pipeline.
Eagle is less a single model than NVIDIA's internal R&D pipeline for multimodal AI, now open-sourced with three generations of VLMs and a grounding specialist.
It gives Python developers on Windows and Linux a fully offline shortcut for detecting faces, extracting landmarks, and scoring similarity.
screenpipe continuously records your screen and audio locally so AI can search, summarize, and act on everything you’ve done without sending data to the cloud.
LichtFeld Studio wraps the entire 3D Gaussian Splatting pipeline—training, editing, exporting, automating—into a single C++ desktop app instead of a chain of Python scripts.
It locally inpaints over hard-coded subtitles and text watermarks in videos and images so you never have to upload frames to a cloud API.
VGGT replaces the traditional multi-stage 3D reconstruction pipeline with a single feed-forward model that predicts cameras, depth, and geometry from one or many images in seconds.
This tool automates image and video annotation by plugging dozens of SOTA models into a single PyQt6 GUI.
Koharu automates the full manga-translation pipeline—detection, OCR, inpainting, and text rendering—entirely on your local machine.





