Computer Vision

Computer Vision

underdogs breaking out
01
baidu/Unlimited-OCR
+23% /wk +611 ★/dayaccelerating

It wants to parse entire documents in one shot without the model getting stuck in repetitive loops.

18.7k Python Computer Vision · explained Feature
02
Robbyant/lingbot-map
+15% /wk +330 ★/daycooling

It reconstructs 3D scenes from streaming video in real time without per-scene optimization, using a feed-forward transformer that remembers trajectory and corrects drift as it goes.

15.3k Python Computer Vision · explained
03
wiltodelta/remove-ai-watermarks
+9.1% /wk +56 ★/dayaccelerating

It removes the visible Gemini sparkle, invisible SynthID fingerprints, and C2PA metadata that AI image generators embed in every output.

4.3k Python Computer Vision · explained
04
SakuraMathcraft/LaTeXSnipper
+8.4% /wk +9.1 ★/daycooling

LaTeXSnipper bundles screenshot OCR, handwriting recognition, symbolic computation, and Office plugins into a single offline desktop app so you can actually use the math you capture.

759 Python Computer Vision · explained
05
aiptimizer/TurboOCR
+5.5% /wk +4.4 ★/daysteady

TurboOCR exists because waiting for a vision-language model to read a receipt is a waste of GPU time.

556 C++ Computer Vision · explained
06
NVlabs/alpamayo
+2.8% /wk +7.8 ★/daysteady

An open 10B-parameter VLA model that predicts driving trajectories while spelling out the causal reasoning behind every turn and lane change.

1.9k Python Domain Apps · explained
07
Tencent-Hunyuan/HunyuanOCR
+2.8% /wk +7.5 ★/daysteady

HunyuanOCR-1.5 speeds up lightweight OCR vision-language models by drafting tokens with a block-diffusion model and closing capability gaps with an agent-driven data pipeline.

1.9k Python Computer Vision · explained
08
NVlabs/alpasim
+2.4% /wk +3.9 ★/daysteady

AlpaSim exists so researchers can validate end-to-end self-driving policies in a closed-loop Python sandbox where renderers, physics, and traffic are swappable microservices.

1.1k Python Agents · explained
09
roboflow/sports
+2.3% /wk +18 ★/dayaccelerating

A testing ground where hard computer vision problems—ball tracking, jersey OCR, player re-ID—get solved with reusable tools.

5.2k Python Computer Vision · explained
10
datalab-to/chandra
+1.8% /wk +30 ★/dayaccelerating

It converts images and PDFs into structured HTML, Markdown, or JSON while reconstructing tables, forms, and handwriting that most OCR tools reduce to plain text soup.

11.8k Python Computer Vision · explained
11
roboflow/rf-detr
+1.6% /wk +20 ★/dayaccelerating

RF-DETR is Roboflow’s bet that a DINOv2 transformer backbone can finally beat YOLO on both speed and accuracy in real-world detection and segmentation tasks.

8.7k Python Computer Vision · explained
12
ankandrew/fast-alpr
+1.4% /wk +1.4 ★/daysteady

Built to let you swap out the plate detector and OCR engine without trashing the rest of the pipeline.

731 Python Computer Vision · explained
13
NVlabs/Eagle
+1.3% /wk +6.0 ★/daysteady

Eagle is less a single model than NVIDIA's internal R&D pipeline for multimodal AI, now open-sourced with three generations of VLMs and a grounding specialist.

3.3k Python Language Models · explained
15
screenpipe/screenpipe
+1.2% /wk +34 ★/daycooling

screenpipe continuously records your screen and audio locally so AI can search, summarize, and act on everything you’ve done without sending data to the cloud.

20.5k Rust Agents · explained
16
MrNeRF/LichtFeld-Studio
+1.0% /wk +4.7 ★/daysteady

LichtFeld Studio wraps the entire 3D Gaussian Splatting pipeline—training, editing, exporting, automating—into a single C++ desktop app instead of a chain of Python scripts.

3.4k C++ Computer Vision · explained
17
YaoFANGUK/video-subtitle-remover
+0.9% /wk +16 ★/daysteady

It locally inpaints over hard-coded subtitles and text watermarks in videos and images so you never have to upload frames to a cloud API.

12k Python Computer Vision · explained
18
facebookresearch/vggt
+0.9% /wk +18 ★/dayaccelerating

VGGT replaces the traditional multi-stage 3D reconstruction pipeline with a single feed-forward model that predicts cameras, depth, and geometry from one or many images in seconds.

14k Python Computer Vision · explained
19
CVHub520/X-AnyLabeling
+0.9% /wk +13 ★/dayaccelerating

This tool automates image and video annotation by plugging dozens of SOTA models into a single PyQt6 GUI.

9.9k Python Data Tooling · explained
20
mayocream/koharu
+0.9% /wk +6.3 ★/dayaccelerating

Koharu automates the full manga-translation pipeline—detection, OCR, inpainting, and text rendering—entirely on your local machine.

4.9k Rust Domain Apps · explained
loading more…

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.