isl-org/DPT
A Vision Transformer architecture for dense prediction tasks like depth estimation and semantic segmentation.

Not currently ranked — collecting fresh signals.
star history
DPT provides pre-trained Vision Transformer models for monocular depth estimation and semantic segmentation. The models use a hybrid architecture combining traditional convolutional layers with transformer encoders, outputting dense pixel-level predictions. The repository includes inference code and downloadable model weights for both tasks.
Frequently asked
- What is isl-org/DPT?
- A Vision Transformer architecture for dense prediction tasks like depth estimation and semantic segmentation.
- Is DPT open source?
- Yes — isl-org/DPT is open source, released under the MIT license.
- What language is DPT written in?
- isl-org/DPT is primarily written in Python.
- How popular is DPT?
- isl-org/DPT has 2.3k stars on GitHub.
- Where can I find DPT?
- isl-org/DPT is on GitHub at https://github.com/isl-org/DPT.