RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual Words
Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu, Feiyue Huang, Rongrong Ji
Abstract
Recent progress on visual question answering has explored the merits of grid features for vision language tasks. Meanwhile, transformer-based models have shown remarkable performance in various sequence prediction problems. However, the spatial information loss of grid features caused by flattening operation, as well as the defect of the transformer model in distinguishing visual words and non visual words, are still left unexplored. In this paper, we first propose Grid-Augmented (GA) module, in which relative geometry features between grids are incorporated to enhance visual representations. Then, we build a BERTbased language model to extract language context and propose Adaptive-Attention (AA) module on top of a transformer decoder to adaptively measure the contribution of visual and language cues before making decisions for word prediction. To prove the generality of our proposals, we apply the two modules to the vanilla transformer model to build our Relationship-Sensitive Transformer (RSTNet) for image captioning task. The proposed model is tested on the MSCOCO benchmark, where it achieves new state-ofart results on both the Karpathy test split and the online test server. Source code is available at GitHub 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8bd8293-aebf-4fbb-bb2d-66db90649d02Cited by top-tier papers37
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie et al.AAAI 2022 · 185 citations
- End-to-End Transformer Based Model for Image CaptioningYiyu Wang, Jungang Xu, Yingfei SunAAAI 2022 · 178 citations
- Injecting Semantic Concepts into End-to-End Image CaptioningZhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lin Liang et al.CVPR 2022 · 125 citations
- DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationBowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu et al.ICLR 2024 · 110 citations
Builds on8
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen et al.AAAI 2021 · 206 citations
- In Defense of Grid Features for Visual Question AnsweringHuaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller et al.CVPR 2020
Related papers
- Normalized and Geometry-Aware Self-Attention Network for Image CaptioningLongteng Guo, Jing Liu, Xinxin Zhu, Peng Yao et al.CVPR 2020
- Improving Intra- and Inter-Modality Visual Relation for Image CaptioningYong Wang, Wenkai Zhang, Qing Liu, Zhengyuan Zhang et al.ACM MM 2020 · 24 citations
- Direction Relation Transformer for Image CaptioningZeliang Song, Xiaofei Zhou, Linhua Dong, Jianlong Tan et al.ACM MM 2021 · 31 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Position-Augmented Transformers with Entity-Aligned Mesh for TextVQAXuanyu Zhang, Qing YangACM MM 2021 · 14 citations
