NVlabs/GCVit
Global Context Vision Transformer (GC ViT) is a PyTorch vision transformer model for image classification, object detection, and semantic segmentation.

Not currently ranked — collecting fresh signals.
star history
GC ViT introduces global context attention mechanisms to vision transformers, enabling efficient capture of long-range dependencies across images. The model achieves competitive performance on ImageNet classification, COCO object detection, and ADE20K semantic segmentation benchmarks. It provides pretrained checkpoints and training code as an official NVIDIA implementation.
Frequently asked
- What is NVlabs/GCVit?
- Global Context Vision Transformer (GC ViT) is a PyTorch vision transformer model for image classification, object detection, and semantic segmentation.
- Is GCVit open source?
- Yes — NVlabs/GCVit is an open-source project tracked on heatdrop.
- What language is GCVit written in?
- NVlabs/GCVit is primarily written in Python.
- How popular is GCVit?
- NVlabs/GCVit has 450 stars on GitHub.
- Where can I find GCVit?
- NVlabs/GCVit is on GitHub at https://github.com/NVlabs/GCVit.