QPrompt-R1: Real-Time Reasoning for Domain-Generalized Semantic Segmentation via Group-Relative Query Alignment
Fengyuan Lu, Zixuan Duan, Xunzhi Xiang, Zhicheng Zhang, Wenbin Li, Yang Gao, Qi Fan
摘要
Deploying semantic segmentation in driving and robotics requires both real-time inference and robustness to domain shifts, formalized as Real-Time Domain-Generalized Semantic Segmentation (RT-DGSS), a challenge not fully addressed. Existing methods treat real-time (RT) inference and domain generalization (DG) separately, with DG improving robustness but lacking real-time performance. To tackle the RT-DGSS problem, we identify that the bottleneck in DG is the prediction head, not the backbone. We introduce QPrompt-R1, a real-time Query-Prompt architecture based on the powerful VFM backbone. QPrompt-R1 integrates reasoning by injecting learnable queries into the final transformer block, leveraging contextual learning to enhance segmentation performance under domain shifts while maintaining real-time inference. To further optimize reasoning without extra inference cost, we introduce a Group Relative Query Alignment (GRQA) training objective, which strengthens the relationship between queries and image tokens through group-relative advantage supervision, unlocking the domain generalization potential of VFMs. QPrompt-R1 achieves 54 FPS, delivering strong performance in synthetic-to-real transfer, real-to-real generalization, and robustness under adverse conditions. GRQA functions as a plug-and-play module, improving DGSS methods such as REIN (+1.2) and SoMA (+0.6) without introducing inference-time overhead. The code is available at QPrompt-R1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
相关 Paper
- ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing ImagesMuhammad Naseer SubhaniCVPR 2026 · 被引用 2 次
- Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic SegmentationZhixiang Wei, Lin Chen, Yi Jin, Xiaoxiao Ma 等CVPR 2024 · 被引用 61 次
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic SegmentationSeogkyu Jeon, Kibeom Hong, Hyeran ByunICCV 2025 · 被引用 2 次
- SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot SegmentationShi-Feng Peng, Guolei Sun, Yong Li, Hongsong Wang 等AAAI 2025 · 被引用 7 次
- LENS: Learning to Segment Anything with Unified Reinforced ReasoningLianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng 等AAAI 2026 · 被引用 7 次
