Learning from Unique Perspectives: User-aware Saliency Modeling
Shi Chen, Nachiappan Valliappan, Shaolei Shen, Xinyu Ye, Kai Kohlhoff, Junfeng He
摘要
Everyone is unique. Given the same visual stimuli, people's attention is driven by both salient visual cues and their own inherent preferences. Knowledge of visual preferences not only facilitates understanding of fine-grained attention patterns of diverse users, but also has the potential of benefiting the development of customized applications. Nevertheless, existing saliency models typically limit their scope to attention as it applies to the general population and ignore the variability between users' behaviors. In this paper, we identify the critical roles of visual preferences in attention modeling, and for the first time study the problem of user-aware saliency modeling. Our work aims to advance attention research from three distinct perspectives: (1) We present a new model with the flexibility to capture attention patterns of various combinations of users, so that we can adaptively predict personalized attention, user group attention, and general saliency at the same time with one single model; (2) To augment models with knowledge about the composition of attention from different users, we further propose a principled learning method to understand visual attention in a progressive manner; and (3) We carry out extensive analyses on publicly available saliency datasets to shed light on the roles of visual preferences. Experimental results on diverse stimuli, including naturalistic images and web pages, demonstrate the advantages of our method in capturing the distinct visual behaviors of different users and the general saliency of visual stimuli.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang 等NeurIPS 2024 · 被引用 43 次
- UniAR: A Unified model for predicting human Attention and Responses on visual contentPeizhao Li, Junfeng He, Gang Li, Rachit Bhargava 等NeurIPS 2024 · 被引用 17 次
- Interpreting Radiologist's Intention from Eye Movements in Chest X-ray DiagnosisTrong-Thang Pham, Anh Nguyen, Zhigang Deng, Carol C. Wu 等ACM MM 2025 · 被引用 1 次
- Attend to Anything: Foundation Model for Unified Human Attention ModelingWenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun ZhaoICML 2026
- CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath ModelingTrong-Thang Pham, Akash Awasthi, Saba Khan, Esteban Duran Marti 等ICCV 2025
它引用的顶会 Paper4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Predicting Visual Importance Across Graphic Design TypesCamilo Fosco, Vincent Casser, Amish Kumar Bedi, Peter O'Donovan 等UIST 2020 · 被引用 55 次
- Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency AnalysisEldon Schoop, Xin Zhou, Gang Li, Zhourong Chen 等CHI 2022 · 被引用 29 次
- Fantastic Answers and Where to Find Them: Immersive Question-Directed Visual AttentionMing Jiang, Shi Chen, Jinhui Yang, Qi ZhaoCVPR 2020
相关 Paper
- Beyond Average: Individualized Visual Scanpath PredictionXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2024
- PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point PredictionHanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao 等ACM MM 2025
- Inferring Attention Shift Ranks of Objects for Image SaliencyAvishek Siris, Jianbo Jiao, Gary K. L. Tam, Xianghua Xie 等CVPR 2020
- Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head AttentionUttaran Bhattacharya, Gang Wu, Stefano Petrangeli, Viswanathan Swaminathan 等ACM MM 2022 · 被引用 5 次
- Multi-view Attentive Variational Learning for Group RecommendationWen Yang, Jiajie Xu, Rui Zhou, Lu Chen 等ICDE 2024 · 被引用 5 次
