InterRVOS: Interaction-Aware Referring Video Object Segmentation
Woojeong Jin, Seongchan Kim, Jaeho Lee, Seungryong Kim
Abstract
Referring video object segmentation (RVOS) aims to segment objects in a video described by a natural language expression. However, most existing approaches focus on segmenting only the referred object (typically the actor), even when the expression clearly describes an interaction involving multiple objects with distinct roles. For instance, "A throwing B" implies a directional interaction, but standard RVOS segments only the actor (A), neglecting other involved target objects (B). In this paper, we introduce Interaction-aware Referring Video Object Segmentation (InterRVOS), a novel task that focuses on the modeling of interactions. It requires the model to segment the actor and target objects separately, reflecting their asymmetric roles in an interaction. This task formulation enables fine-grained understanding of object relationships, as many video events are defined by such relationships rather than individual objects. To support this task, we propose a new evaluation protocol that separately evaluates actor and target segmentation, enabling more accurate assessment of the model's ability to distinguish and segment actor and target roles. We also present InterRVOS-127K, a large-scale dataset with over 127K automatically annotated expressions, including interaction expressions annotated with distinct masks for actor and target objects. Furthermore, we develop ReVIOSa, an MLLM-based architecture that introduces interaction-aware special tokens and leverages an attention mask loss to enhance role-specific segmentation. Extensive experiments show that ReVIOSa not only outperforms existing baselines on our proposed InterRVOS-127K evaluation set, but also achieves strong performance on standard RVOS benchmarks. Our project page is available at: https://cvlab-kaist.github.io/InterRVOS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2fb45c2-1c5e-4c85-9a24-1924270d4072Cited by top-tier papers1
Ask how each one uses itBuilds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 359 citations
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 242 citations
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 150 citations
Related papers
- RVAS: Referring Video Active Exploration and SegmentationHengrui Hu, Weiwei Gao, Zipei Zhang, Henghui DingICML 2026
- DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object SegmentationWenxuan Cheng, Ming Dai, Huimin Lu, Wankou YangCVPR 2026
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual SegmentationKaining Ying, Henghui Ding, Guangquan Jie, Yu-Gang JiangICCV 2025 · 3 citations
- Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object SegmentationTianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan et al.CVPR 2026 · 8 citations
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded SegmentationAli Athar, Xueqing Deng, Liang-Chieh ChenCVPR 2025
