Personalized Image Descriptions from Attention Sequences
Ruoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal, Abe Leite, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras
摘要
People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description generation focus on linguistic style alone, with no prior work leveraging individual viewing patterns. We address this gap by explicitly modeling personalized viewing behavior as a core factor in description generation. Our method, DEPER (DEscription-PERception persona encoder), learns a subject embedding that captures both linguistic style and viewing behavior, guided by an auxiliary attention-prediction task. A lightweight adapter aligns these embeddings with a frozen vision-language model, enabling few-shot personalization without retraining. Across four datasets spanning diverse viewing tasks and both short and detailed descriptions, DEPER achieves a 24% average improvement, showing that modeling personalized attention produces more human-aligned and high-quality descriptions. We posit that understanding how people see helps predict what they say; modeling human diversity in perception can improve both performance and human alignment in multi-modal systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 被引用 94 次
- Human Attention in Image Captioning: Dataset and AnalysisSen He, Hamed Rezazadegan Tavakoli, Ali Borji, Nicolas PugeaultICCV 2019 · 被引用 55 次
相关 Paper
- Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized SentencesDingyi Yang, Hongyu Chen, Xinglin Hou, Tiezheng Ge 等ACM MM 2023 · 被引用 5 次
- Nested Attention: Semantic-aware Attention Values for Concept PersonalizationOr Patashnik, Rinon Gal, Daniil Ostashev, Sergey Tulyakov 等SIGGRAPH 2025 · 被引用 6 次
- Language Does Matter for Cross-Domain Few-Shot Visual Feature EnhancementFei Zhou, Xiwen Zhang, Qingqing Qiu, Lei Zhang 等CVPR 2026
- Who You Are Decides How You TellShuang Wu, Shaojing Fan, Zhiqi Shen, Mohan S. Kankanhalli 等ACM MM 2020 · 被引用 2 次
- Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image CustomizationYeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin 等AAAI 2025 · 被引用 6 次
