Improving Panoptic Narrative Grounding by Harnessing Semantic Relationships and Visual Confirmation
Tianyu Guo, Haowei Wang, Yiwei Ma, Jiayi Ji, Xiaoshuai Sun
摘要
Recent advancements in single-stage Panoptic Narrative Grounding (PNG) have demonstrated significant potential. These methods predict pixel-level masks by directly matching pixels and phrases. However, they often neglect the modeling of semantic and visual relationships between phrase-level instances, limiting their ability for complex multi-modal reasoning in PNG. To tackle this issue, we propose XPNG, a “differentiation-refinement-localization” reasoning paradigm for accurately locating instances or regions. In XPNG, we introduce a Semantic Context Convolution (SCC) module to leverage semantic priors for generating distinctive features. This well-crafted module employs a combination of dynamic channel-wise convolution and pixel-wise convolution to embed semantic information and establish inter-object relationships guided by semantics. Subsequently, we propose a Visual Context Verification (VCV) module to provide visual cues, eliminating potential space biases introduced by semantics and further refining the visual features generated by the previous module. Extensive experiments on PNG benchmark datasets reveal that our approach achieves state-of-the-art performance, significantly outperforming existing methods by a considerable margin and yielding a 3.9-point improvement in overall metrics. Our codes and results are available at our project webpage: https://github.com/TianyuGoGO/XPNG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- F-LMM: Grounding Frozen Large Multimodal ModelsSize Wu, Sheng Jin, Wenwei Zhang, Lumin Xu 等CVPR 2025
- ACL: Activating Capability of Linear Attention for Image RestorationYubin Gu, Yuan Meng, Jiayi Ji, Xiaoshuai SunCVPR 2025
它引用的顶会 Paper25
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- K-Net: Towards Unified Image SegmentationWenwei Zhang, Jiangmiao Pang, Kai Chen, Chen Change LoyNeurIPS 2021 · 被引用 500 次
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 被引用 359 次
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen 等CVPR 2022 · 被引用 319 次
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan 等ACM MM 2022 · 被引用 314 次
相关 Paper
- Towards Real-Time Panoptic Narrative Grounding by an End-to-End Grounding NetworkHaowei Wang, Jiayi Ji, Yiyi Zhou, Yongjian Wu 等AAAI 2023 · 被引用 18 次
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang 等ACM MM 2022 · 被引用 12 次
- Panoptic Narrative GroundingCristina González, Nicolás Ayobi, Isabela Hernández, José Hernández 等ICCV 2021 · 被引用 30 次
- Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative GroundingHongyu Li, Tianrui Hui, Zihan Ding, Jing Zhang 等ACM MM 2024 · 被引用 2 次
- Semi-Supervised Panoptic Narrative GroundingDanni Yang, Jiayi Ji, Xiaoshuai Sun, Haowei Wang 等ACM MM 2023 · 被引用 7 次
