YuqingWang1029/VisTR
End-to-end video instance segmentation framework using transformer architecture.

Not currently ranked — collecting fresh signals.
star history
VisTR implements an end-to-end approach to video instance segmentation by applying transformer architecture to jointly process video frames and predict instance masks across time. The model leverages a transformer-based detection framework (DETR) adapted for video understanding, enabling unified instance tracking and segmentation without additional post-processing. It is designed for video understanding tasks in computer vision research.
Frequently asked
- What is YuqingWang1029/VisTR?
- End-to-end video instance segmentation framework using transformer architecture.
- Is VisTR open source?
- Yes — YuqingWang1029/VisTR is open source, released under the Apache-2.0 license.
- What language is VisTR written in?
- YuqingWang1029/VisTR is primarily written in Python.
- How popular is VisTR?
- YuqingWang1029/VisTR has 757 stars on GitHub.
- Where can I find VisTR?
- YuqingWang1029/VisTR is on GitHub at https://github.com/YuqingWang1029/VisTR.