Context-Aware Group Captioning via Self-Attention and Contrastive Features
Zhuowan Li, Quan Tran, Long Mai, Zhe Lin, Alan L. Yuille
Abstract
While image captioning has progressed rapidly, existing works focus mainly on describing single images. In this paper, we introduce a new task, context-aware group captioning, which aims to describe a group of target images in the context of another group of related reference images. Context-aware group captioning requires not only summarizing information from both the target and reference image group but also contrasting between them. To solve this problem, we propose a framework combining selfattention mechanism with contrastive feature construction to effectively summarize common information from each image group while capturing discriminative information between them. To build the dataset for this task, we propose to group the images and generate the group captions based on single image captions using scene graphs matching. Our datasets are constructed on top of the public Conceptual Captions dataset and our new Stock Captions dataset. Experiments on the two datasets show the effectiveness of our method on this new task. 1 * This work has been done during the first author's internship at Adobe. 1 Related Datasets and code are released at https://lizw14. github.io/project/groupcap .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3c895dc-c7fe-4580-b7ab-bce8bf973c08Cited by top-tier papers10
- AutoAttend: Automated Attention Representation SearchChaoyu Guan, Xin Wang, Wenwu ZhuICML 2021 · 46 citations
- Rethinking the Reference-based Distinctive Image CaptioningYangjun Mao, Long Chen, Zhihong Jiang, Dong Zhang et al.ACM MM 2022 · 22 citations
- Group-based Distinctive Image Captioning with Memory AttentionJiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. ChanACM MM 2021 · 20 citations
- Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real ImagesZhuowan Li, Elias Stengel-Eskin, Yixiao Zhang, Cihang Xie et al.ICCV 2021 · 19 citations
- Exploring Group Video Captioning with Efficient Relational ApproximationWang Lin, Tao Jin, Ye Wang, Wenwen Pan et al.ICCV 2023 · 17 citations
Builds on4
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang et al.ICCV 2019 · 694 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Neural Architecture Search for Lightweight Non-Local NetworksYingwei Li, Xiaojie Jin, Jieru Mei, Xiaochen Lian et al.CVPR 2020
Related papers
- Object Relation Attention for Image Paragraph CaptioningLi-Chuan Yang, Chih-Yuan Yang, Jane Yung-jen HsuAAAI 2021 · 17 citations
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 160 citations
- Question-controlled Text-aware Image CaptioningAnwen Hu, Shizhe Chen, Qin JinACM MM 2021 · 11 citations
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha et al.ACL 2024
- Multi-Perspective Video CaptioningYi Bin, Xindi Shang, Bo Peng, Yujuan Ding et al.ACM MM 2021 · 14 citations
