Discriminative Perception via Anchored Description for Reasoning Segmentation
Tao Yang, Qing Zhou, Yanliang Li, Qi Wang
Abstract
Reasoning segmentation increasingly employs reinforcement learning to generate explanatory reasoning chains that guide Multimodal Large Language Models. While these geometric rewards are primarily confined to guiding the final localization, they are incapable of discriminating whether the reasoning process remains anchored on the referred region or strays into irrelevant context. Lacking this discriminative guidance, the model's reasoning often devolves into unfocused and verbose chains that ultimately fail to disambiguate and perceive the target in complex scenes. This suggests a need to complement the RL objective with Discriminative Perception, an ability to actively distinguish a target from its context. To realize this, we propose DPAD to compel the model to generate a descriptive caption of the referred object, which is then used to explicitly discriminate by contrasting the caption's semantic relevance to the referred object against the wider context. By optimizing for this discriminative capability, the model is forced to focus on the unique attributes of the target, leading to a more converged and efficient reasoning chain. The descriptive caption also serves as an interpretability rationale that aligns with the segmentation. Experiments on the benchmarks confirm the validity of our approach, delivering substantial performance gains, with the cIoU on Rea-sonSeg increasing by 3.09% and the reasoning chain length decreasing by approximately 42%. Code is available at https://github.com/mrazhou/DPAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 990641ce-5d04-4da4-aa48-3a7b4188060bBuilds on12
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-TuningYibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang et al.NeurIPS 2025 · 102 citations
- LogicSeg: Parsing Visual Semantics with Neural Logic Learning and ReasoningLiulei Li, Wenguan Wang, Yang YiICCV 2023 · 52 citations
- PixelLM: Pixel Reasoning with Large Multimodal ModelZhongwei Ren, Zhicheng Huang, Yunchao Wei, Yao Zhao et al.CVPR 2024 · 48 citations
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningJiaqi Huang, Zunnan Xu, Jun Zhou, Ting Liu et al.NeurIPS 2025 · 33 citations
Related papers
- RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-ThoughtYi Lu, Jiawang Cao, Yongliang Wu, Bozheng Li et al.ACL 2025 · 15 citations
- Don't Overthink with Pixels: Efficient Reasoning for SegmentationSong Wang, Gongfan Fang, Lingdong Kong, Xiangtai Li et al.ICML 2026
- RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided SegmentationXingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li et al.ICML 2026 · 3 citations
- DRSeg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language ModelsYulin He, Wei Chen, Zhikang Jian, Tianhang Guo et al.ICML 2026
- CG-Reasoner: Centroid-Guided Positional Reasoning Segmentation for Medical Imaging with a Robust Visual-Text Consistency MetricLakshmikar R., Ming MaCVPR 2026
