ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
Tianming Liang, Kun-Yu Lin, Chaolei Tan, Jianguo Zhang, Wei-Shi Zheng, Jian-Fang Hu
Abstract
Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existing methods still exhibit a noticeable gap when considering all these aspects. In this work, we propose ReferDINO, a strong RVOS model that inherits region-level vision-language alignment from foundational visual grounding models, and is further endowed with pixel-level dense perception and cross-modal spatiotemporal reasoning. In detail, ReferDINO integrates two key components: 1) a grounding-guided deformable mask decoder that utilizes location prediction to progressively guide mask prediction through differentiable deformation mechanisms; 2) an object-consistent temporal enhancer that injects pretrained time-varying text features into inter-frame interaction to capture object-aware dynamic changes. Moreover, a confidence-aware query pruning strategy is designed to accelerate object decoding without compromising model performance. Extensive experimental results on five benchmarks demonstrate that our ReferDINO significantly outperforms previous methods (e.g., +3.9% (J&F) on Ref-YouTube-VOS) with real-time inference speed (51 FPS).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d71c6a1e-8809-4f99-b254-d7b783893f00Cited by top-tier papers5
- Deforming Videos to Masks: Flow Matching for Referring Video SegmentationZanyi Wang, Dengyang Jiang, Liuzhuozheng Li, Sizhe Dang et al.ICLR 2026 · 10 citations
- MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression SegmentationChangli Wu, Haodong Wang, Jiayi Ji, Yutian Yao et al.CVPR 2026 · 8 citations
- Panoptic Captioning: An Equivalence Bridge for Image and TextKun-Yu Lin, Hongjun Wang, Weining Ren, Kai HanNeurIPS 2025 · 7 citations
- Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object MemoryZhengtong Zhu, Jiaqing Fan, Zhixuan Liu, Fanzhang LiAAAI 2026 · 1 citation
- Matting Anything 2: Towards Video Matting for AnythingChenyi Zhang, Yiheng Lin, Yunchao Wei, Hongsong Wang et al.ICLR 2026
Builds on26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world DetectionLewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang et al.NeurIPS 2022 · 285 citations
Related papers
- Multi-Level Representation Learning with Semantic Alignment for Referring Video Object SegmentationDongming Wu, Xingping Dong, Ling Shao, Jianbing ShenCVPR 2022 · 55 citations
- VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal SegmentationJihwan Hong, Jaeyoung DoCVPR 2026 · 2 citations
- DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object SegmentationWenxuan Cheng, Ming Dai, Huimin Lu, Wankou YangCVPR 2026
- OnlineRefer: A Simple Online Baseline for Referring Video Object SegmentationDongming Wu, Tiancai Wang, Yuang Zhang, Xiangyu Zhang et al.ICCV 2023 · 82 citations
- LoSh: Long-Short Text Joint Prediction Network for Referring Video Object SegmentationLinfeng Yuan, Miaojing Shi, Zijie Yue, Qijun ChenCVPR 2024 · 12 citations
