Bottom-Up and Bidirectional Alignment for Referring Expression Comprehension
Liuwu Li, Yuqi Bu, Yi Cai
摘要
In this paper, we propose a one-stage approach to improve referring expression comprehension (REC) which aims at grounding the referent according to a natural language expression. We observe that humans understand referring expressions through a fine-to-coarse bottom-up way, and bidirectionally obtain vision-language information between image and text. Inspired by this, we define the language granularity and the vision granularity. Otherwise, existing methods do not follow the mentioned way of human understanding in referring expression. Motivated by our observation and to address the limitations of existing methods, we propose a bottom-up and bidirectional alignment (BBA) framework. Our method constructs the cross-modal alignment starting from fine-grained representation to coarse-grained representation and bidirectionally obtains vision-language information between image and text. Based on the structure of BBA, we further propose a progressive visual attribute decomposing approach to decompose visual proposals into several independent spaces to enhance the bottom-up alignment framework. Experiments on five benchmark datasets of RefCOCO, RefCOCO+, ReferItGame, RefCOCOg and Flick30K show that our approach obtains +2.16%, +4.47%, +2.85%, +3.44%, and +2.91% improvements over the one-stage SOTA approaches, which validates the effectiveness of our approach.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- Open-Vocabulary Object Detection via Scene Graph DiscoveryHengcan Shi, Munawar Hayat, Jianfei CaiACM MM 2023 · 被引用 20 次
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang 等ACM MM 2022 · 被引用 12 次
- RefCrowd: Grounding the Target in Crowd with Referring ExpressionsHeqian Qiu, Hongliang Li, Taijin Zhao, Lanxiao Wang 等ACM MM 2022 · 被引用 10 次
- Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression ComprehensionYaxian Wang, Henghui Ding, Shuting He, Xudong Jiang 等AAAI 2025 · 被引用 9 次
- Semi-Supervised Panoptic Narrative GroundingDanni Yang, Jiayi Ji, Xiaoshuai Sun, Haowei Wang 等ACM MM 2023 · 被引用 7 次
相关 Paper
- Revisiting Counterfactual Problems in Referring Expression ComprehensionZhihan Yu, Ruifan LiCVPR 2024 · 被引用 6 次
- Co-Grounding Networks With Semantic Attention for Referring Expression Comprehension in VideosSijie Song, Xudong Lin, Jiaying Liu, Zongming Guo 等CVPR 2021
- Towards Further Comprehension on Referring Expression with RationaleRengang Li, Baoyu Fan, Xiaochuan Li, Runze Zhang 等ACM MM 2022 · 被引用 2 次
- Exploring Logical Reasoning for Referring Expression ComprehensionYing Cheng, Ruize Wang, Jiashuo Yu, Rui-Wei Zhao 等ACM MM 2021 · 被引用 12 次
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu 等CVPR 2025
