It exists to turn a single photo into metric depth, camera pose, and a 3D point cloud using nothing but a native binary—no Python stack required.
Computer Vision
underdogs breaking outIt gives Python developers on Windows and Linux a fully offline shortcut for detecting faces, extracting landmarks, and scoring similarity.
The Autoware Foundation wants to ship production-grade ADAS code you can actually audit, train, and modify.
TurboOCR exists because waiting for a vision-language model to read a receipt is a waste of GPU time.
It removes the visible Gemini sparkle, invisible SynthID fingerprints, and C2PA metadata that AI image generators embed in every output.
LaTeXSnipper bundles screenshot OCR, handwriting recognition, symbolic computation, and Office plugins into a single offline desktop app so you can actually use the math you capture.
It reconstructs 3D scenes from streaming video in real time without per-scene optimization, using a feed-forward transformer that remembers trajectory and corrects drift as it goes.
A curated index of hands-on machine learning, generative AI, and data-science projects, with about a third marked as end-to-end builds.
Eagle is less a single model than NVIDIA's internal R&D pipeline for multimodal AI, now open-sourced with three generations of VLMs and a grounding specialist.
Koharu automates the full manga-translation pipeline—detection, OCR, inpainting, and text rendering—entirely on your local machine.
AlpaSim exists so researchers can validate end-to-end self-driving policies in a closed-loop Python sandbox where renderers, physics, and traffic are swappable microservices.
UniFace wraps two dozen specialized face models into one Python library so you can detect, recognize, track, parse, and anonymize faces without managing a zoo of dependencies.
It converts images and PDFs into structured HTML, Markdown, or JSON while reconstructing tables, forms, and handwriting that most OCR tools reduce to plain text soup.
An ONNX-exported, multi-engine OCR toolkit that runs offline on basically anything.
It corrals dozens of optical flow architectures into one PyTorch Lightning framework so you can train and benchmark them without maintaining forty separate codebases.
BetterGI automates Genshin Impact by reading the screen and simulating clicks, not by modifying game memory.
RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.
HunyuanOCR-1.5 speeds up lightweight OCR vision-language models by drafting tokens with a block-diffusion model and closing capability gaps with an agent-driven data pipeline.
An open 10B-parameter VLA model that predicts driving trajectories while spelling out the causal reasoning behind every turn and lane change.
screenpipe continuously records your screen and audio locally so AI can search, summarize, and act on everything you’ve done without sending data to the cloud.







