4DAnyone generates dense multi-view video from a single casual clip, feeding standard 4D Gaussian Splatting pipelines without a real camera rig.
Computer Vision
underdogs breaking outA Flask dashboard and Android client that repurpose e-waste into a privacy-first computer-vision pipeline with structured event logging.
It packages a C++17 video analytics runtime, a browser-based pipeline editor, and async VLM nodes into a single deployable appliance stack for edge hardware.
This plugin wires an upstream vision toolkit into DeepSeek Harness so text-only models can inspect images, locate UI elements, and diff screenshots without swapping LLMs.
It gives Python developers on Windows and Linux a fully offline shortcut for detecting faces, extracting landmarks, and scoring similarity.
MOSS-VL is an open-weight 11B-parameter model series built to understand video in real time, deciding on its own when to respond and when to keep observing.
A transformer-based object detector that claims to beat YOLO at its own real-time game, with official Paddle and PyTorch implementations.
RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.
ByteTrack is a real-time multi-object tracker that rescues low-confidence bounding boxes instead of discarding them, recovering occluded objects and reducing fragmented trajectories.
UniFace wraps two dozen specialized face models into one Python library so you can detect, recognize, track, parse, and anonymize faces without managing a zoo of dependencies.
It removes the visible Gemini sparkle, invisible SynthID fingerprints, and C2PA metadata that AI image generators embed in every output.
CompreFace wraps FaceNet and InsightFace in a Dockerized REST API so developers can add face recognition, verification, and detection to apps without building ML pipelines.
Unity sample scenes that wire Meta Quest's passthrough camera into object detection, AI queries, WebRTC streams, and spatial shaders.
A self-hosted NVR that runs object detection, face recognition, and motion alerts entirely on local hardware.
A self-hosted NVR that runs YOLOv9 detection on RTSP streams locally and pings your phone with AI-summarized alerts.
A curated index of hands-on machine learning, generative AI, and data-science projects, with about a third marked as end-to-end builds.
An ONNX-exported, multi-engine OCR toolkit that runs offline on basically anything.
It locally inpaints over hard-coded subtitles and text watermarks in videos and images so you never have to upload frames to a cloud API.
A Qt-based image annotation tool that added AI assistance without becoming a SaaS product.
It estimates temporally consistent depth for arbitrarily long monocular videos by caching temporal attention states, trading diffusion bloat for speed.




