A Flask dashboard and Android client that repurpose e-waste into a privacy-first computer-vision pipeline with structured event logging.
Computer Vision
underdogs · picking up speedIt packages a C++17 video analytics runtime, a browser-based pipeline editor, and async VLM nodes into a single deployable appliance stack for edge hardware.
This plugin wires an upstream vision toolkit into DeepSeek Harness so text-only models can inspect images, locate UI elements, and diff screenshots without swapping LLMs.
A transformer-based object detector that claims to beat YOLO at its own real-time game, with official Paddle and PyTorch implementations.
ByteTrack is a real-time multi-object tracker that rescues low-confidence bounding boxes instead of discarding them, recovering occluded objects and reducing fragmented trajectories.
RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.
CompreFace wraps FaceNet and InsightFace in a Dockerized REST API so developers can add face recognition, verification, and detection to apps without building ML pipelines.
A Qt-based image annotation tool that added AI assistance without becoming a SaaS product.
A vision model that counts anything you can describe, from cattle to cancer cells, by placing a point on every instance.
A Python tool that uses DeepSeek OCR to convert scanned academic books into structured Markdown or EPUB without calling LLM APIs.
A free Android image editor that packs dozens of tools—from basic crops to OCR, EXIF editing, and AI upscaling—into one open-source package.
Saber-Translator runs locally as a web app, detects speech bubbles, OCRs Japanese text, translates via your choice of AI provider, then inpaints the original text and renders new Chinese text back in.
It reconstructs 3D scenes from streaming video in real time without per-scene optimization, using a feed-forward transformer that remembers trajectory and corrects drift as it goes.
A self-hosted NVR that runs object detection, face recognition, and motion alerts entirely on local hardware.
UNet++ redesigns skip connections into dense, nested pathways so medical image segmentation no longer depends on guessing the perfect U-Net depth.
SAM 3 exists so you can segment and track objects in images and video by describing them with text, points, or boxes—no custom training required.
FiftyOne exists because model performance is usually a data problem, and staring at raw image folders doesn't scale.
It locally inpaints over hard-coded subtitles and text watermarks in videos and images so you never have to upload frames to a cloud API.
CVAT is the self-hosted annotation workbench that lets teams draw bounding boxes, polygons, and 3D cuboids on their own hardware while plugging in their own models to do the boring parts.
An ONNX-exported, multi-engine OCR toolkit that runs offline on basically anything.