Temporal Collection and Distribution for Referring Video Object Segmentation
Jiajin Tang, Ge Zheng, Sibei Yang
Abstract
Referring video object segmentation aims to segment a referent throughout a video sequence according to a natural language expression. It requires aligning the natural language expression with the objects’ motions and their dynamic associations at the global video level but segmenting objects at the frame level. To achieve this goal, we propose to simultaneously maintain a global referent token and a sequence of object queries, where the former is responsible for capturing video-level referent according to the language expression, while the latter serves to better locate and segment objects with each frame. Furthermore, to explicitly capture object motions and spatial-temporal cross-modal reasoning over objects, we propose a novel temporal collection-distribution mechanism for interacting between the global referent token and object queries. Specifically, the temporal collection mechanism collects global information for the referent token from object queries to the temporal motions to the language expression. In turn, the temporal distribution first distributes the referent token to the referent sequence across all frames and then performs efficient cross-frame reasoning between the referent sequence and object queries in every frame. Experimental results show that our method outperforms state-of-the-art methods on all benchmarks consistently and significantly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e0bba2b-de22-4fff-87ef-73d298ead373Cited by top-tier papers23
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language ModelsGe Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou et al.NeurIPS 2023 · 252 citations
- Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object SegmentationShaofei Huang, Rui Ling, Hongyu Li, Tianrui Hui et al.AAAI 2025 · 24 citations
- LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationLinfeng Yuan, Miaojing Shi, Zijie Yue, Qijun ChenCVPR 2024 · 12 citations
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment FormatsJiaye Qian, Ge Zheng, Yuchen Zhu, Sibei YangNeurIPS 2025 · 11 citations
- The Devil is in the Object Boundary: Towards Annotation-free Instance Segmentation using Foundation ModelsCheng Shi, Sibei YangICLR 2024 · 10 citations
Builds on31
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 359 citations
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
Related papers
- Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video GroundingZihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin et al.CVPR 2023
- Mitigating Query Selection Bias in Referring Video Object SegmentationDingwei Zhang, Dong Zhang, Jinhui TangACM MM 2025 · 1 citation
- SOC: Semantic-Assisted Object Cluster for Referring Video Object SegmentationZhuoyan Luo, Yicheng Xiao, Yong Liu, Shuyan Li et al.NeurIPS 2023 · 89 citations
- Look Around Before Locating: Considering Content and Structure Information for Visual GroundingShiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He et al.AAAI 2025 · 3 citations
- VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal SegmentationJihwan Hong, Jaeyoung DoCVPR 2026 · 2 citations
