PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative Grounding
Zihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang, Xiaoming Wei, Xiaolin Wei, Si Liu
摘要
Panoptic Narrative Grounding (PNG) is an emerging task whose goal is to segment visual objects of things and stuff categories described by dense narrative captions of a still image. The previous two-stage approach first extracts segmentation region proposals by an off-the-shelf panoptic segmentation model, then conducts coarse region-phrase matching to ground the candidate regions for each noun phrase. However, the two-stage pipeline usually suffers from the performance limitation of low-quality proposals in the first stage and the loss of spatial details caused by region feature pooling, as well as complicated strategies designed for things and stuff categories separately. To alleviate these drawbacks, we propose a one-stage end-to-end Pixel-Phrase Matching Network (PPMN), which directly matches each phrase to its corresponding pixels instead of region proposals and outputs panoptic segmentation by simple combination. Thus, our model can exploit sufficient and finer cross-modal semantic correspondence from the supervision of densely annotated pixel-phrase pairs rather than sparse region-phrase pairs. In addition, we also propose a Language-Compatible Pixel Aggregation (LCPA) module to further enhance the discriminative ability of phrase features through multi-round refinement, which selects the most compatible pixels for each phrase to adaptively aggregate the corresponding visual context. Extensive experiments show that our method achieves new state-of-the-art performance on the PNG benchmark with 4.0 absolute Average Recall gains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- 3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression SegmentationChangli Wu, Yiwei Ma, Qi Chen, Haowei Wang 等AAAI 2024 · 被引用 40 次
- Semi-Supervised Panoptic Narrative GroundingDanni Yang, Jiayi Ji, Xiaoshuai Sun, Haowei Wang 等ACM MM 2023 · 被引用 7 次
- Improving Panoptic Narrative Grounding by Harnessing Semantic Relationships and Visual ConfirmationTianyu Guo, Haowei Wang, Yiwei Ma, Jiayi Ji 等AAAI 2024 · 被引用 5 次
- Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative GroundingHongyu Li, Tianrui Hui, Zihan Ding, Jing Zhang 等ACM MM 2024 · 被引用 2 次
- 3D-DRES: Detailed 3D Referring Expression SegmentationQi Chen, Changli Wu, Jiayi Ji, Yiwei Ma 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper32
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- K-Net: Towards Unified Image SegmentationWenwei Zhang, Jiangmiao Pang, Kai Chen, Chen Change LoyNeurIPS 2021 · 被引用 500 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
相关 Paper
- Towards Real-Time Panoptic Narrative Grounding by an End-to-End Grounding NetworkHaowei Wang, Jiayi Ji, Yiyi Zhou, Yongjian Wu 等AAAI 2023 · 被引用 18 次
- Panoptic Narrative GroundingCristina González, Nicolás Ayobi, Isabela Hernández, José Hernández 等ICCV 2021 · 被引用 30 次
- LPSNet: A Lightweight Solution for Fast Panoptic SegmentationWeixiang Hong, Qingpei Guo, Wei Zhang, Jingdong Chen 等CVPR 2021
- Dense Video Object Captioning from Disjoint SupervisionXingyi Zhou, Anurag Arnab, Chen Sun, Cordelia SchmidICLR 2025
- TextPSG: Panoptic Scene Graph Generation from Textual DescriptionsChengyang Zhao, Yikang Shen, Zhenfang Chen, Mingyu Ding 等ICCV 2023 · 被引用 24 次
