Guideline-Consistent Segmentation via Multi-Agent Refinement
Vanshika Vats, Ashwani Rathee, James Davis
Abstract
Semantic segmentation in real-world applications often requires not only accurate masks but also strict adherence to textual labeling guidelines. These guidelines are typically complex and long, and both human and automated labeling often fail to follow them faithfully. Traditional approaches depend on expensive task-specific retraining that must be repeated as the guidelines evolve. Although recent open-vocabulary segmentation methods excel with simple prompts, they often fail when confronted with sets of paragraph-length guidelines that specify intricate segmentation rules. To address this, we introduce a multi-agent, training-free framework that coordinates general-purpose vision-language models within an iterative Worker-Supervisor refinement architecture. The Worker performs the segmentation, the Supervisor critiques it against the retrieved guidelines, and a lightweight reinforcement learning stop policy decides when to terminate the loop, ensuring guideline-consistent masks while balancing resource use. Evaluated on the Waymo and ReasonSeg datasets, our method notably outperforms state-of-the-art baselines, demonstrating strong generalization and instruction adherence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided SegmentationXingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li et al.ICML 2026 · 3 citations
- SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning SegmentationTianyuan Du, Haopeng Li, Zhen Fan, Jiarui Zhang et al.ICLR 2026
- LENS: Learning to Segment Anything with Unified Reinforced ReasoningLianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng et al.AAAI 2026 · 7 citations
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian SplattingChangyue Shi, Minghao Chen, Yiping Mao, Chuxiao Yang et al.CVPR 2026 · 8 citations
