Computer Vision

Computer Vision

underdogs · picking up speed
02
cosmo-wander-ai/cosmo-edge
+18% /wk +20 ★/dayaccelerating

It packages a C++17 video analytics runtime, a browser-based pipeline editor, and async VLM nodes into a single deployable appliance stack for edge hardware.

797 C Inference · Serving · explained
03
Anionex/dsh-vision-toolkit
+9.6% /wk +12 ★/dayaccelerating

This plugin wires an upstream vision toolkit into DeepSeek Harness so text-only models can inspect images, locate UI elements, and diff screenshots without swapping LLMs.

872 TypeScript Agents · explained
04
lyuwenyu/RT-DETR
+3.2% /wk +25 ★/dayaccelerating

A transformer-based object detector that claims to beat YOLO at its own real-time game, with official Paddle and PyTorch implementations.

5.5k Python Computer Vision · explained
05
FoundationVision/ByteTrack
+2.4% /wk +23 ★/dayaccelerating

ByteTrack is a real-time multi-object tracker that rescues low-confidence bounding boxes instead of discarding them, recovering occluded objects and reducing fragmented trajectories.

6.7k Python Computer Vision · explained
06
roboflow/rf-detr
+3.0% /wk +40 ★/dayaccelerating

RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.

9.4k Python Computer Vision · explained
07
exadel-inc/CompreFace
+1.8% /wk +22 ★/dayaccelerating

CompreFace wraps FaceNet and InsightFace in a Dockerized REST API so developers can add face recognition, verification, and detection to apps without building ML pipelines.

8.3k Java Computer Vision · explained
08
wkentaro/labelme
+1.0% /wk +24 ★/dayaccelerating

A Qt-based image annotation tool that added AI assistance without becoming a SaaS product.

16.2k Python Data Tooling · explained
09
Mengqi-Lei/count-anything
+0.9% /wk +0.7 ★/daysteady

A vision model that counts anything you can describe, from cattle to cancer cells, by placing a point on every instance.

541 Python Computer Vision · explained
10
oomol-lab/pdf-craft
+0.6% /wk +5.3 ★/dayaccelerating

A Python tool that uses DeepSeek OCR to convert scanned academic books into structured Markdown or EPUB without calling LLM APIs.

6.3k Python Computer Vision · explained
11
T8RIN/ImageToolbox
+0.8% /wk +16 ★/dayaccelerating

A free Android image editor that packs dozens of tools—from basic crops to OCR, EXIF editing, and AI upscaling—into one open-source package.

14.6k Kotlin Computer Vision · explained
12
MashiroSaber03/Saber-Translator
+0.7% /wk +3.7 ★/daysteady

Saber-Translator runs locally as a web app, detects speech bubbles, OCRs Japanese text, translates via your choice of AI provider, then inpaints the original text and renders new Chinese text back in.

3.6k Python Domain Apps · explained
13
Robbyant/lingbot-map
+0.7% /wk +16 ★/dayaccelerating

It reconstructs 3D scenes from streaming video in real time without per-scene optimization, using a feed-forward transformer that remembers trajectory and corrects drift as it goes.

16.9k Python Computer Vision · explained
14
roflcoopter/viseron
+1.2% /wk +6.1 ★/daysteady

A self-hosted NVR that runs object detection, face recognition, and motion alerts entirely on local hardware.

3.5k Python Computer Vision · explained
15
MrGiovanni/UNetPlusPlus
+0.1% /wk +0.3 ★/daysteady

UNet++ redesigns skip connections into dense, nested pathways so medical image segmentation no longer depends on guessing the perfect U-Net depth.

2.7k Python Computer Vision · explained
16
facebookresearch/sam3
+0.6% /wk +10 ★/dayaccelerating

SAM 3 exists so you can segment and track objects in images and video by describing them with text, points, or boxes—no custom training required.

11.6k Python Computer Vision · explained
17
voxel51/fiftyone
+0.2% /wk +2.4 ★/daysteady

FiftyOne exists because model performance is usually a data problem, and staring at raw image folders doesn't scale.

11.1k TypeScript Data Tooling · explained
18
YaoFANGUK/video-subtitle-remover
+1.1% /wk +20 ★/daysteady

It locally inpaints over hard-coded subtitles and text watermarks in videos and images so you never have to upload frames to a cloud API.

12.8k Python Computer Vision · explained
19
cvat-ai/cvat
+0.2% /wk +5.1 ★/daysteady

CVAT is the self-hosted annotation workbench that lets teams draw bounding boxes, polygons, and 3D cuboids on their own hardware while plugging in their own models to do the boring parts.

16.7k Python Data Tooling · explained
20
RapidAI/RapidOCR
+1.1% /wk +12 ★/daysteady

An ONNX-exported, multi-engine OCR toolkit that runs offline on basically anything.

7.8k Python Computer Vision · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.