Context-Aware Group Captioning via Self-Attention and Contrastive Features
Zhuowan Li, Quan Tran, Long Mai, Zhe Lin, Alan L. Yuille
摘要
While image captioning has progressed rapidly, existing works focus mainly on describing single images. In this paper, we introduce a new task, context-aware group captioning, which aims to describe a group of target images in the context of another group of related reference images. Context-aware group captioning requires not only summarizing information from both the target and reference image group but also contrasting between them. To solve this problem, we propose a framework combining selfattention mechanism with contrastive feature construction to effectively summarize common information from each image group while capturing discriminative information between them. To build the dataset for this task, we propose to group the images and generate the group captions based on single image captions using scene graphs matching. Our datasets are constructed on top of the public Conceptual Captions dataset and our new Stock Captions dataset. Experiments on the two datasets show the effectiveness of our method on this new task. 1 * This work has been done during the first author's internship at Adobe. 1 Related Datasets and code are released at https://lizw14. github.io/project/groupcap .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- AutoAttend: Automated Attention Representation SearchChaoyu Guan, Xin Wang, Wenwu ZhuICML 2021 · 被引用 46 次
- Rethinking the Reference-based Distinctive Image CaptioningYangjun Mao, Long Chen, Zhihong Jiang, Dong Zhang 等ACM MM 2022 · 被引用 22 次
- Group-based Distinctive Image Captioning with Memory AttentionJiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. ChanACM MM 2021 · 被引用 20 次
- Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real ImagesZhuowan Li, Elias Stengel-Eskin, Yixiao Zhang, Cihang Xie 等ICCV 2021 · 被引用 19 次
- Exploring Group Video Captioning with Efficient Relational ApproximationWang Lin, Tao Jin, Ye Wang, Wenwen Pan 等ICCV 2023 · 被引用 17 次
它引用的顶会 Paper4
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang 等ICCV 2019 · 被引用 694 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
- Neural Architecture Search for Lightweight Non-Local NetworksYingwei Li, Xiaojie Jin, Jieru Mei, Xiaochen Lian 等CVPR 2020
相关 Paper
- Object Relation Attention for Image Paragraph CaptioningLi-Chuan Yang, Chih-Yuan Yang, Jane Yung-jen HsuAAAI 2021 · 被引用 17 次
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 被引用 160 次
- Question-controlled Text-aware Image CaptioningAnwen Hu, Shizhe Chen, Qin JinACM MM 2021 · 被引用 11 次
- Context-aware Difference Distilling for Multi-change CaptioningYunbin Tu, Liang Li, Li Su, Zheng-Jun Zha 等ACL 2024
- Multi-Perspective Video CaptioningYi Bin, Xindi Shang, Bo Peng, Yujuan Ding 等ACM MM 2021 · 被引用 14 次
