← all repositories

NVlabs/describe-anything

A large multimodal model that generates detailed captions for arbitrary regions of images or video frames.

describe-anything
Not currently ranked — collecting fresh signals.
star history

Describe Anything Model (DAM) takes region annotations (points, boxes, scribbles, masks) on images or video frames and outputs detailed textual descriptions of those regions. For videos, a single frame annotation suffices. The project includes a new evaluation benchmark (DLC-Bench) to assess models on the detailed localized captioning task.

Frequently asked

What is NVlabs/describe-anything?
A large multimodal model that generates detailed captions for arbitrary regions of images or video frames.
Is describe-anything open source?
Yes — NVlabs/describe-anything is open source, released under the Apache-2.0 license.
What language is describe-anything written in?
NVlabs/describe-anything is primarily written in Python.
How popular is describe-anything?
NVlabs/describe-anything has 1.5k stars on GitHub.
Where can I find describe-anything?
NVlabs/describe-anything is on GitHub at https://github.com/NVlabs/describe-anything.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.