Distributed Attention for Grounded Image Captioning
Nenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang, Zhiwen Lin, Yuqiang Ren, Haolei Yuan, Xiaowei Guo, Feiyue Huang, Wenping Wang
Abstract
We study the problem of weakly supervised grounded image captioning. That is, given an image, the goal is to automatically generate a sentence describing the context of the image with each noun word grounded to the corresponding region in the image. This task is challenging due to the lack of explicit fine-grained region word alignments as supervision. Previous weakly supervised methods mainly explore various kinds of regularization schemes to improve attention accuracy. However, their performances are still far from the fully supervised ones. One main issue that has been ignored is that the attention for generating visually groundable words may only focus on the most discriminate parts and can not cover the whole object. To this end, we propose a simple yet effective method to alleviate the issue, termed as partial grounding problem in our paper. Specifically, we design a distributed attention mechanism to enforce the network to aggregate information from multiple spatially different regions with consistent semantics while generating the words. Therefore, the union of the focused region proposals should form a visual region that encloses the object of interest completely. Extensive experiments have demonstrated the superiority of our proposed method compared with the state-of-the-arts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2af28171-3685-4bc2-8985-75b363d7ba6cCited by top-tier papers2
- You Can even Annotate Text with Voice: Transcription-only-Supervised Text SpottingJingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma et al.ACM MM 2022 · 22 citations
- AttentionShift: Iteratively Estimated Part-Based Attention Map for Pointly Supervised Instance SegmentationMingxiang Liao, Zonghao Guo, Yuze Wang, Peng Yuan et al.CVPR 2023
Builds on12
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 · 260 citations
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 100 citations
- Prophet Attention: Predicting Attention with Future AttentionFenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge et al.NeurIPS 2020 · 52 citations
Related papers
- Weakly-Supervised Generation and Grounding of Visual Descriptions with Conditional Generative ModelsEffrosyni Mavroudi, René VidalCVPR 2022 · 5 citations
- More Grounded Image Captioning by Distilling Image-Text Matching ModelYuanen Zhou, Meng Wang, Daqing Liu, Zhenzhen Hu et al.CVPR 2020
- Comprehensive Visual Grounding for Video DescriptionWenhui Jiang, Yibo Cheng, Linxin Liu, Yuming Fang et al.AAAI 2024 · 5 citations
- Diffusion-Assisted Progressive Learning for Weakly Supervised Phrase LocalizationPengyue Lin, Yanyang Hu, Xinjing Liu, Wenqi Jia et al.AAAI 2026
- Detector-Free Weakly Supervised Grounding by SeparationAssaf Arbelle, Sivan Doveh, Amit Alfassy, Joseph Shtok et al.ICCV 2021 · 31 citations
