Lune

ACM MM2021Top-tier venue

Bottom-Up and Bidirectional Alignment for Referring Expression Comprehension

Liuwu Li, Yuqi Bu, Yi Cai

2021Year
11Citations
8Top-tier citations

Abstract

In this paper, we propose a one-stage approach to improve referring expression comprehension (REC) which aims at grounding the referent according to a natural language expression. We observe that humans understand referring expressions through a fine-to-coarse bottom-up way, and bidirectionally obtain vision-language information between image and text. Inspired by this, we define the language granularity and the vision granularity. Otherwise, existing methods do not follow the mentioned way of human understanding in referring expression. Motivated by our observation and to address the limitations of existing methods, we propose a bottom-up and bidirectional alignment (BBA) framework. Our method constructs the cross-modal alignment starting from fine-grained representation to coarse-grained representation and bidirectionally obtains vision-language information between image and text. Based on the structure of BBA, we further propose a progressive visual attribute decomposing approach to decompose visual proposals into several independent spaces to enhance the bottom-up alignment framework. Experiments on five benchmark datasets of RefCOCO, RefCOCO+, ReferItGame, RefCOCOg and Flick30K show that our approach obtains +2.16%, +4.47%, +2.85%, +3.44%, and +2.91% improvements over the one-stage SOTA approaches, which validates the effectiveness of our approach.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 1a8145e8-b701-4bb2-b77d-ef3fcfefbbae

Cited by top-tier papers8

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines