Direction Relation Transformer for Image Captioning
Zeliang Song, Xiaofei Zhou, Linhua Dong, Jianlong Tan, Li Guo
摘要
Image captioning is a challenging task that combines computer vision and natural language processing for generating a textual description of the content within an image. Recently, Transformer-based encoder-decoder architectures have shown great success in image captioning, where multi-head attention mechanism is utilized to capture the contextual interactions between object regions. However, such methods regard region features as a bag of tokens without considering the directional relationships between them, making it hard to understand the relative position between objects in the image and generate correct captions effectively. In this paper, we propose a novel Direction Relation Transformer to improve the orientation perception between visual features by incorporating the relative direction embedding into multi-head attention, termed DRT. We first generate the relative direction matrix according to the positional information of the object regions, and then explore three forms of direction-aware multi-head attention to integrate the direction embedding into Transformer architecture. We conduct experiments on challenging Microsoft COCO image captioning benchmark. The quantitative and qualitative results demonstrate that, by integrating the relative directional relation, our proposed approach achieves significant improvements over all evaluation metrics compared with baseline model, e.g., DRT improves task-specific metric CIDEr score from 129.7% to 133.2% on the offline ''Karpathy'' test split.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Progressive Tree-Structured Prototype Network for End-to-End Image CaptioningPengpeng Zeng, Jinkuan Zhu, Jingkuan Song, Lianli GaoACM MM 2022 · 被引用 35 次
- Image Captioning with Multi-Context Synthetic DataFeipeng Ma, Yizhou Zhou, Fengyun Rao, Yueyi Zhang 等AAAI 2024 · 被引用 22 次
相关 Paper
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao 等AAAI 2021 · 被引用 349 次
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 被引用 4 次
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen 等ICCV 2019 · 被引用 107 次
- Image Captioning with Context-Aware Auxiliary GuidanceZeliang Song, Xiaofei Zhou, Zhendong Mao, Jianlong TanAAAI 2021 · 被引用 36 次
- Multimodal Attention with Image Text Spatial Relationship for OCR-Based Image CaptioningJing Wang, Jinhui Tang, Jiebo LuoACM MM 2020 · 被引用 55 次
