Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
Haichao Jiang, Tianming Liang, Wei-Shi Zheng, Jian-Fang Hu
Abstract
Referring Video Object Segmentation (RVOS) aims to segment objects in videos based on textual queries. Current methods mainly rely on large-scale supervised fine-tuning (SFT) of Multi-modal Large Language Models (MLLMs). However, this paradigm suffers from heavy data dependence and limited scalability against the rapid evolution of MLLMs. Although recent zero-shot approaches offer a flexible alternative, their performance remains significantly behind SFT-based methods, due to the straightforward workflow designs. To address these limitations, we propose Refer-Agent, a collaborative multi-agent system with alternating reasoning-reflection mechanisms. This system decomposes RVOS into step-by-step reasoning process. During reasoning, we introduce a Coarse-to-Fine frame selection strategy to ensure the frame diversity and textual relevance, along with a Dynamic Focus Layout that adaptively adjusts the agent's visual focus. Furthermore, we propose a Chain-of-Reflection mechanism, which employs a Questioner-Responder pair to generate a self-reflection chain, enabling the system to verify intermediate results and generates feedback for next-round reasoning refinement. Extensive experiments on five challenging benchmarks demonstrate that Refer-Agent significantly outperforms state-of-the-art methods, including both SFT-based models and zero-shot approaches. Moreover, Refer-Agent is flexible and enables fast integration of new MLLMs without any additional fine-tuning costs. Code will be released at https : / / github . com / iSEE -Laboratory / Refer-Agent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6c3cfb6-66a9-49f3-a9d8-3729b118c98aBuilds on25
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion ExpressionsHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 242 citations
- End-to-End Referring Video Object Segmentation with Multimodal TransformersAdam Botach, Evgenii Zheltonozhskii, Chaim BaskinCVPR 2022 · 150 citations
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang et al.NeurIPS 2024 · 147 citations
Related papers
- RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided SegmentationXingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li et al.ICML 2026 · 3 citations
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai et al.AAAI 2026
- CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for VideosShiu-Hong Kao, Yu-Wing Tai, Chi-Keung TangICLR 2026 · 8 citations
- Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual SegmentationJinxing Zhou, Yanghao Zhou, Mingfei Han, Tong Wang et al.AAAI 2026 · 5 citations
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video ReasoningYang Ding, Xin Lai, Yizhen Zhang, Wei Li et al.ICLR 2026 · 26 citations
