Say As You Wish: Fine-Grained Control of Image Caption Generation With Abstract Scene Graphs
Shizhe Chen, Qin Jin, Peng Wang, Qi Wu
Abstract
Humans are able to describe image contents with coarse to fine details as they wish. However, most image captioning models are intention-agnostic which can not generate diverse descriptions according to different user intentions initiatively. In this work, we propose the Abstract Scene Graph (ASG) structure to represent user intention in fine-grained level and control what and how detailed the generated description should be. The ASG is a directed graph consisting of three types of abstract nodes (object, attribute, relationship) grounded in the image without any concrete semantic labels. Thus it is easy to obtain either manually or automatically. From the ASG, we propose a novel ASG2Caption model, which is able to recognise user intentions and semantics in the graph, and therefore generate desired captions according to the graph structure. Our model achieves better controllability conditioning on ASGs than carefully designed baselines on both VisualGenome and MSCOCO datasets. It also significantly improves the caption diversity via automatically sampling diverse ASGs as control signals.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c115c64-57ff-46a4-8a4f-3c69e3ae390dCited by top-tier papers42
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image CaptioningJun Chen, Han Guo, Kai Yi, Boyang Li et al.CVPR 2022 · 169 citations
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu et al.CVPR 2022 · 116 citations
- Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingJiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang et al.CVPR 2022 · 89 citations
- Tailor: Generating and Perturbing Text with Semantic ControlsAlexis Ross, Tongshuang Wu, Hao Peng, Matthew E. Peters et al.ACL 2022 · 85 citations
- PPDL: Predicate Probability Distribution based Loss for Unbiased Scene Graph GenerationWei Li, Haiwei Zhang, Qijie Bai, Guoqing Zhao et al.CVPR 2022 · 64 citations
Builds on3
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Sequential Latent Spaces for Modeling the Intention During Diverse Image CaptioningJyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2019 · 71 citations
Related papers
- SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioningXu Zhang, Jin Yuan, Hanwang Zhang, Guojin Zhong et al.AAAI 2025 · 2 citations
- In Defense of Scene Graphs for Image CaptioningKien Nguyen, Subarna Tripathi, Bang Du, Tanaya Guha et al.ICCV 2021 · 55 citations
- Topic Scene Graph Generation by Attention Distillation from CaptionWenbin Wang, Ruiping Wang, Xilin ChenICCV 2021 · 16 citations
- Unconditional Scene Graph GenerationSarthak Garg, Helisa Dhamo, Azade Farshad, Sabrina Musatian et al.ICCV 2021 · 30 citations
- Towards Accurate Text-Based Image Captioning With Content Diversity ExplorationGuanghui Xu, Shuaicheng Niu, Mingkui Tan, Yucheng Luo et al.CVPR 2021
