Fantastic Answers and Where to Find Them: Immersive Question-Directed Visual Attention
Ming Jiang, Shi Chen, Jinhui Yang, Qi Zhao
Abstract
While most visual attention studies focus on bottom-up attention with restricted field-of-view, real-life situations are filled with embodied vision tasks. The role of attention is more significant in the latter due to the information overload, and attention to the most important regions is critical to the success of tasks. The effects of visual attention on task performance in this context have also been widely ignored. This research addresses a number of challenges to bridge this research gap, on both the data and model aspects. Specifically, we introduce the first dataset of top-down attention in immersive scenes. The Immersive Questiondirected Visual Attention (IQVA) dataset features visual attention and corresponding task performance (i.e., answer correctness). It consists of 975 questions and answers collected from people viewing 360°videos in a head-mounted display. Analyses of the data demonstrate a significant correlation between people's task performance and their eye movements, suggesting the role of attention in task performance. With that, a neural network is developed to encode the differences of correct and incorrect attention and jointly predict the two. The proposed attention model for the first time takes into account answer correctness, whose outputs naturally distinguish important regions from distractions. This study with new data and features may enable new tasks that leverage attention and answer correctness, and inspire new research that reveals the process behind decision making in performing various tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Query and Attention Augmentation for Knowledge-Based Explainable ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2022 · 14 citations
- RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360deg Image Quality AssessmentYujia Wang, Yuyan Li, Jiuming Liu, Fang-Lue Zhang et al.CVPR 2026 · 3 citations
- Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual ForagingBo Wang, Dingwei Tan, Yen-Ling Kuo, Zhaowei Sun et al.CVPR 2025
- Learning from Unique Perspectives: User-aware Saliency ModelingShi Chen, Nachiappan Valliappan, Shaolei Shen, Xinyu Ye et al.CVPR 2023
- Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using CapsulesAisha Urooj Khan, Hilde Kuehne, Kevin Duarte, Chuang Gan et al.CVPR 2021
Related papers
- Quantification of Users' Visual Attention During Everyday Mobile Device InteractionsMihai Bâce, Sander Staal, Andreas BullingCHI 2020 · 25 citations
- FixationNet: Forecasting Eye Fixations in Task-Oriented Virtual EnvironmentsZhiming Hu, Andreas Bulling, Sheng Li, Guoping WangIEEE VR 2021 · 76 citations
- Atari-HEAD: Atari Human Eye-Tracking and Demonstration DatasetRuohan Zhang, Calen Walshe, Zhuode Liu, Lin Guan et al.AAAI 2020 · 77 citations
- Comparison of Visual Saliency for Dynamic Point Clouds: Task-free vs. Task-dependentXuemei Zhou, Irene Viola, Silvia Rossi, Pablo CésarIEEE VR 2025 · 5 citations
- Predicting Human Scanpaths in Visual Question AnsweringXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2021
