Entangled Transformer for Image Captioning
Guang Li, Linchao Zhu, Ping Liu, Yi Yang
摘要
In image captioning, the typical attention mechanisms are arduous to identify the equivalent visual signals especially when predicting highly abstract words. This phenomenon is known as the semantic gap between vision and language. This problem can be overcome by providing semantic attributes that are homologous to language. Thanks to the inherent recurrent nature and gated operating mechanism, Recurrent Neural Network (RNN) and its variants are the dominating architectures in image captioning. However, when designing elaborate attention mechanisms to integrate visual inputs and semantic attributes, RNN-like variants become unflexible due to their complexities. In this paper, we investigate a Transformer-based sequence modeling framework, built only with attention layers and feedforward layers. To bridge the semantic gap, we introduce EnTangled Attention (ETA) that enables the Transformer to exploit semantic and visual information simultaneously. Furthermore, Gated Bilateral Controller (GBC) is proposed to guide the interactions between the multimodal information. We name our model as ETA-Transformer. Remarkably, ETA-Transformer achieves state-of-the-art performance on the MSCOCO image captioning dataset. The ablation studies validate the improvements of our proposed modules. 𝑇 𝑣 : a bunch of fruit sitting in a sink. 𝑇 𝑠 : a table with a lot of food on it. 𝐸𝑇𝐴 : a bowl of fruits and vegetables on a stove. 𝑇 𝑣 : a baby girl laying on a bed holding a toy. 𝑇 𝑠 : a baby girl laying on a bed with a bed. 𝐸𝑇𝐴: a baby sitting on a bed with a bottle. 𝑇 𝑣 : a giraffe eating from a feeder in a zoo. 𝑇 𝑠 : a giraffe eating a tree with a tree in background. 𝐸𝑇𝐴: a giraffe eating hay out of a feeder. 𝑇 𝑣 : a clock hanging from a wall next to a window. 𝑇 𝑠 : a large clock sitting on top of a wall. 𝐸𝑇𝐴: a clock hanging on the side of a building.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- DeepSVG: A Hierarchical Generative Network for Vector Graphics AnimationAlexandre Carlier, Martin Danelljan, Alexandre Alahi, Radu TimofteNeurIPS 2020 · 被引用 247 次
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen 等AAAI 2021 · 被引用 206 次
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image CaptioningJun Chen, Han Guo, Kai Yi, Boyang Li 等CVPR 2022 · 被引用 169 次
相关 Paper
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 被引用 4 次
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
- Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningWeizhi Nie, Jiesi Li, Ning Xu, An-An Liu 等ACM MM 2021 · 被引用 9 次
- End-to-End Transformer Based Model for Image CaptioningYiyu Wang, Jungang Xu, Yingfei SunAAAI 2022 · 被引用 178 次
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 被引用 124 次
