DIFNet: Boosting Visual Information Flow for Image Captioning
Mingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou, Chao Chen, Jiaxin Gu, Xing Sun, Rongrong Ji
Abstract
Current Image Captioning (IC) methods predict textual words sequentially based on the input visual information from the visual feature extractor and the partially generated sentence information. However, for most cases, the partially generated sentence may dominate the target word prediction due to the insufficiency of visual information, making the generated descriptions irrelevant to the content of the given image. In this paper, we propose a Dual Information Flow Network (DIFNet <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> Source code is available at: https://github.com/mrwu-mac/DIFNet) to address this issue, which takes segmentation feature as another visual information source to enhance the contribution of visual information for prediction. To maximize the use of two information flows, we also propose an effective feature fusion module termed Iterative Independent Layer Normalization (IILN) which can condense the most relevant inputs while retraining modality-specific information in each flow. Experiments show that our method is able to enhance the dependence of prediction on visual information, making word prediction more focused on the visual content, and thus achieves new state-of-the-art performance on the MSCOCO dataset, e.g., 136.2 CIDEr on COCO Karpathy test split.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 048d8ba4-7f99-4f21-8d3d-af2419e595efCited by top-tier papers14
- DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationBowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu et al.ICLR 2024 · 110 citations
- ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language ModelsMingrui Wu, Xinyue Cai, Jiayi Ji, Jiale Li et al.NeurIPS 2024 · 50 citations
- With a Little Help from your own Past: Prototypical Memory Networks for Image CaptioningManuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi et al.ICCV 2023 · 33 citations
- Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language ModelsMingrui Wu, Jiayi Ji, Oucheng Huang, Jiale Li et al.ICML 2024 · 32 citations
- Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image CaptioningZhiyue Liu, Jinyuan Liu, Fanrong MaAAAI 2024 · 23 citations
Builds on10
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- Hierarchy Parsing for Image CaptioningTing Yao, Yingwei Pan, Yehao Li, Tao MeiICCV 2019 · 183 citations
- Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer FusionYikai Wang, Fuchun Sun, Ming Lu, Anbang YaoACM MM 2020 · 66 citations
Related papers
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen et al.ICCV 2019 · 107 citations
- Distilled Cross-Combination Transformer for Image Captioning with Dual Refined Visual FeaturesJunbo Hu, Zhixin LiACM MM 2024 · 8 citations
- Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image CaptioningXinzhi Dong, Chengjiang Long, Wenju Xu, Chunxia XiaoACM MM 2021 · 75 citations
- Semi-Autoregressive Image CaptioningXu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang et al.ACM MM 2021 · 22 citations
- Show, Edit and Tell: A Framework for Editing Image CaptionsFawaz Sammani, Luke Melas-KyriaziCVPR 2020
