Tube-Link: A Flexible Cross Tube Framework for Universal Video Segmentation
Xiangtai Li, Haobo Yuan, Wenwei Zhang, Guangliang Cheng, Jiangmiao Pang, Chen Change Loy
Abstract
Video segmentation aims to segment and track every pixel in diverse scenarios accurately. In this paper, we present Tube-Link, a versatile framework that addresses multiple core tasks of video segmentation with a unified architecture. Our framework is a near-online approach that takes a short subclip as input and outputs the corresponding spatial-temporal tube masks. To enhance the modeling of cross-tube relationships, we propose an effective way to perform tube-level linking via attention along the queries. In addition, we introduce temporal contrastive learning to instance-wise discriminative features for tubelevel association. Our approach offers flexibility and efficiency for both short and long video inputs, as the length of each subclip can be varied according to the needs of datasets or scenarios. Tube-Link outperforms existing specialized architectures by a significant margin on five video segmentation datasets. Specifically, it achieves almost 13% relative improvements on VIPSeg and 4% improvements on KITTI-STEP over the strong baseline Video K-Net. When using a ResNet50 backbone on Youtube-VIS-2019 and 2021, Tube-Link boosts IDOL by 3% and 4%, respectively. Code is available at https://github. com/lxtGH/Tube-Link .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36d0dfc8-66c6-4210-854b-14cd03fb9922Cited by top-tier papers6
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationThong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen et al.AAAI 2025 · 8 citations
- CAVIS: Context-Aware Video Instance SegmentationSeunghun Lee, Jiwan Seo, Kiljoon Han, Minwoo Choi et al.ICCV 2025 · 4 citations
- Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsYikang Zhou, Tao Zhang, Shilin Xu, Shihao Chen et al.ICCV 2025 · 2 citations
- SAM 2: Segment Anything in Images and VideosNikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu et al.ICLR 2025
- Exploiting Temporal State Space Sharing for Video Semantic SegmentationSyed Ariff Syed Hesham, Yun Liu, Guolei Sun, Henghui Ding et al.CVPR 2025
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- K-Net: Towards Unified Image SegmentationWenwei Zhang, Jiangmiao Pang, Kai Chen, Chen Change LoyNeurIPS 2021 · 500 citations
Related papers
- Video K-Net: A Simple, Strong, and Unified Baseline for Video SegmentationXiangtai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen et al.CVPR 2022 · 71 citations
- TubeFormer-DeepLab: Video Mask TransformerDahun Kim, Jun Xie, Huiyu Wang, Siyuan Qiao et al.CVPR 2022 · 28 citations
- MinVIS: A Minimal Video Instance Segmentation Framework without Video-based TrainingDe-An Huang, Zhiding Yu, Anima AnandkumarNeurIPS 2022 · 135 citations
- TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training ModelBo Pang, Yizhuo Li, Yifan Zhang, Muchen Li et al.CVPR 2020
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li et al.ICCV 2021 · 124 citations
