Tube-Link: A Flexible Cross Tube Framework for Universal Video Segmentation
Xiangtai Li, Haobo Yuan, Wenwei Zhang, Guangliang Cheng, Jiangmiao Pang, Chen Change Loy
摘要
Video segmentation aims to segment and track every pixel in diverse scenarios accurately. In this paper, we present Tube-Link, a versatile framework that addresses multiple core tasks of video segmentation with a unified architecture. Our framework is a near-online approach that takes a short subclip as input and outputs the corresponding spatial-temporal tube masks. To enhance the modeling of cross-tube relationships, we propose an effective way to perform tube-level linking via attention along the queries. In addition, we introduce temporal contrastive learning to instance-wise discriminative features for tubelevel association. Our approach offers flexibility and efficiency for both short and long video inputs, as the length of each subclip can be varied according to the needs of datasets or scenarios. Tube-Link outperforms existing specialized architectures by a significant margin on five video segmentation datasets. Specifically, it achieves almost 13% relative improvements on VIPSeg and 4% improvements on KITTI-STEP over the strong baseline Video K-Net. When using a ResNet50 backbone on Youtube-VIS-2019 and 2021, Tube-Link boosts IDOL by 3% and 4%, respectively. Code is available at https://github. com/lxtGH/Tube-Link .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationThong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen 等AAAI 2025 · 被引用 8 次
- CAVIS: Context-Aware Video Instance SegmentationSeunghun Lee, Jiwan Seo, Kiljoon Han, Minwoo Choi 等ICCV 2025 · 被引用 4 次
- Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsYikang Zhou, Tao Zhang, Shilin Xu, Shihao Chen 等ICCV 2025 · 被引用 2 次
- SAM 2: Segment Anything in Images and VideosNikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu 等ICLR 2025
- Exploiting Temporal State Space Sharing for Video Semantic SegmentationSyed Ariff Syed Hesham, Yun Liu, Guolei Sun, Henghui Ding 等CVPR 2025
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
- K-Net: Towards Unified Image SegmentationWenwei Zhang, Jiangmiao Pang, Kai Chen, Chen Change LoyNeurIPS 2021 · 被引用 500 次
相关 Paper
- Video K-Net: A Simple, Strong, and Unified Baseline for Video SegmentationXiangtai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen 等CVPR 2022 · 被引用 71 次
- TubeFormer-DeepLab: Video Mask TransformerDahun Kim, Jun Xie, Huiyu Wang, Siyuan Qiao 等CVPR 2022 · 被引用 28 次
- MinVIS: A Minimal Video Instance Segmentation Framework without Video-based TrainingDe-An Huang, Zhiding Yu, Anima AnandkumarNeurIPS 2022 · 被引用 135 次
- TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training ModelBo Pang, Yizhuo Li, Yifan Zhang, Muchen Li 等CVPR 2020
- Crossover Learning for Fast Online Video Instance SegmentationShusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li 等ICCV 2021 · 被引用 124 次
