Who You Are Decides How You Tell
Shuang Wu, Shaojing Fan, Zhiqi Shen, Mohan S. Kankanhalli, Anthony K. H. Tung
摘要
Image captioning is gaining significance in multiple applications such as content-based visual search and chat-bots. Much of the recent progress in this field embraces a data-driven approach without deep consideration of human behavioural characteristics. In this paper, we focus on human-centered automatic image captioning. Our study is based on the intuition that different people will generate a variety of image captions for the same scene, as their knowledge and opinion about the scene may differ. In particular, we first perform a series of human studies to investigate what influences human description of a visual scene. We identify three main factors: a person's knowledge level of the scene, opinion on the scene, and gender. Based on our human study findings, we propose a novel human-centered algorithm that is able to generate human-like image captions. We evaluate the proposed model through traditional evaluation metrics, diversity metrics, and human-based evaluation. Experimental results demonstrate the superiority of our proposed model on generating diverse human-like image captions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Towards Accurate Text-Based Image Captioning With Content Diversity ExplorationGuanghui Xu, Shuaicheng Niu, Mingkui Tan, Yucheng Luo 等CVPR 2021
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal 等CVPR 2026 · 被引用 2 次
- Generating Diverse and Descriptive Image Captions Using Visual ParaphrasesLixin Liu, Jiajun Tang, Xiaojun Wan, Zongming GuoICCV 2019 · 被引用 48 次
- Tell as You Want: Customizing Image Narrative with Knowledge and ThoughtsZiwei Yao, Qian Wang, Ruiping Wang, Xilin ChenAAAI 2026 · 被引用 1 次
- IC3: Image Captioning by Committee ConsensusDavid Chan, Austin Myers, Sudheendra Vijayanarasimhan, David A. Ross 等EMNLP 2023 · 被引用 8 次
