Because commodity WiFi already bounces off your body, RuView uses cheap ESP32 nodes to detect presence, vital signs, and even body pose without cameras or wearables.
Computer Vision
big names · picking up speedOmniParser turns raw screenshots into structured, labeled UI elements so vision-language models can finally click what they mean to click.
A Tencent research project that restores degraded faces by tapping into the rich priors locked inside a pretrained StyleGAN2 model.
CMU's real-time multi-person pose estimator detects body, face, hands, and feet simultaneously—body runtime stays flat even as the crowd grows.
It renders high-quality novel views of real-world scenes at 30 fps by replacing costly neural radiance fields with optimized 3D Gaussians.
An OCR model that asks how few vision tokens an LLM needs before it can no longer read the page.
It exists to let developers run customized vision, text, and audio machine learning across mobile, web, and edge hardware without cloud round-trips.
YOLOv5 made real-time object detection as easy as `torch.hub.load`, then exported to everything from iOS to edge chips.
To give developers a zero-shot image segmentation model that generates masks from a click or a bounding box, no retraining required.
It turns face swaps and lip-syncs into queued, retryable batch jobs instead of one-off scripts.
DeepFace wraps a zoo of pre-trained face models into a single Python API so you can verify identities, search databases, and analyze attributes without hand-rolling a Keras pipeline.
Real-ESRGAN turns the ESRGAN research model into a practical tool for upscaling and restoring real-world images and videos using only synthetic training data.
Google Research releases all its code and datasets in one place, which has grown so large that the README treats a full clone as a hazard.
Ultralytics wants to stop you from stitching together separate repos for every computer vision task by bundling detection, segmentation, tracking, and pose estimation into one YOLO-backed package.
It glues together CRAFT detection and CRNN recognition so you can pull text out of images without tuning neural networks yourself.
It bundles detection, recognition, alignment, and reconstruction into a single research-grade toolbox.
It turns images and PDFs into structured JSON and Markdown so your RAG pipeline doesn't have to squint.
It rounds up hundreds of AI project links so you don't have to hunt them down yourself.
OpenCV is an open-source C++ computer vision library whose own README acts as a portal rather than a product page.
Frigate performs real-time, local object detection on IP camera streams using OpenCV and TensorFlow, designed to integrate tightly with Home Assistant.


