LQMFormer: Language-Aware Query Mask Transformer for Referring Image Segmentation
Nisarg A. Shah, Vibashan VS, Vishal M. Patel
摘要
Referring Image Segmentation (RIS) aims to segment objects from an image based on a language description. Recent advancements have introduced transformer-based methods that leverage cross-modal dependencies, significantly enhancing performance in referring segmentation tasks. These methods are designed such that each query predicts different masks. However, RIS inherently requires a single-mask prediction, leading to a phenomenon known as Query Collapse, where all queries yield the same mask prediction. This reduces the generalization capability of the RIS model for complex or novel scenarios. To address this issue, we propose a Multi-modal Query Feature Fusion technique, characterized by two innovative designs: (1) Gaussian enhanced Multi-Modal Fusion, a novel visual grounding mechanism that enhances overall representation by extracting rich local visual information and global visual-linguistic relationships, and (2) A Dynamic Query Module that produces a diverse set of queries through a scoring network where the network selectively focuses on queries for objects referred to in the language description. Moreover, we show that including an auxiliary loss to increase the distance between mask representations of different queries further enhances performance and mitigates query collapse. Extensive experiments conducted on four benchmark datasets validate the effectiveness of our framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word EmphasisYuji Wang, Jingchen Ni, Yong Liu, Chun Yuan 等AAAI 2025 · 被引用 23 次
- IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression SegmentationQi Chen, Changli Wu, Jiayi Ji, Yiwei Ma 等AAAI 2025 · 被引用 5 次
- Interaction-aware Representation Modeling With Co-Occurrence Consistency for Egocentric Hand-Object ParsingYUEJIAO SU, Yi Wang, Lei Yao, Yawen Cui 等ICLR 2026 · 被引用 5 次
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image SegmentationZhenjie Mao, Yuhuan Yang, Chaofan Ma, Dongsheng Jiang 等NeurIPS 2025 · 被引用 4 次
- GenMask: Adapting DiT for Segmentation via Direct Mask GenerationYuhuan Yang, Xianwei Zhuang, Yuxuan Cai, Chaofan Ma 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 被引用 270 次
- Mask Grounding for Referring Image SegmentationYong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu 等CVPR 2024
- Improving Target Presence and Plurality Recognition for Generalized Referring Image SegmentationNamyup Kim, Jinsung Lee, Suha KwakAAAI 2026
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen 等CVPR 2022 · 被引用 319 次
