Fantastic Answers and Where to Find Them: Immersive Question-Directed Visual Attention
Ming Jiang, Shi Chen, Jinhui Yang, Qi Zhao
摘要
While most visual attention studies focus on bottom-up attention with restricted field-of-view, real-life situations are filled with embodied vision tasks. The role of attention is more significant in the latter due to the information overload, and attention to the most important regions is critical to the success of tasks. The effects of visual attention on task performance in this context have also been widely ignored. This research addresses a number of challenges to bridge this research gap, on both the data and model aspects. Specifically, we introduce the first dataset of top-down attention in immersive scenes. The Immersive Questiondirected Visual Attention (IQVA) dataset features visual attention and corresponding task performance (i.e., answer correctness). It consists of 975 questions and answers collected from people viewing 360°videos in a head-mounted display. Analyses of the data demonstrate a significant correlation between people's task performance and their eye movements, suggesting the role of attention in task performance. With that, a neural network is developed to encode the differences of correct and incorrect attention and jointly predict the two. The proposed attention model for the first time takes into account answer correctness, whose outputs naturally distinguish important regions from distractions. This study with new data and features may enable new tasks that leverage attention and answer correctness, and inspire new research that reveals the process behind decision making in performing various tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Query and Attention Augmentation for Knowledge-Based Explainable ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2022 · 被引用 14 次
- RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360deg Image Quality AssessmentYujia Wang, Yuyan Li, Jiuming Liu, Fang-Lue Zhang 等CVPR 2026 · 被引用 3 次
- Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual ForagingBo Wang, Dingwei Tan, Yen-Ling Kuo, Zhaowei Sun 等CVPR 2025
- Learning from Unique Perspectives: User-aware Saliency ModelingShi Chen, Nachiappan Valliappan, Shaolei Shen, Xinyu Ye 等CVPR 2023
- Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using CapsulesAisha Urooj Khan, Hilde Kuehne, Kevin Duarte, Chuang Gan 等CVPR 2021
相关 Paper
- Quantification of Users' Visual Attention During Everyday Mobile Device InteractionsMihai Bâce, Sander Staal, Andreas BullingCHI 2020 · 被引用 25 次
- FixationNet: Forecasting Eye Fixations in Task-Oriented Virtual EnvironmentsZhiming Hu, Andreas Bulling, Sheng Li, Guoping WangIEEE VR 2021 · 被引用 76 次
- Atari-HEAD: Atari Human Eye-Tracking and Demonstration DatasetRuohan Zhang, Calen Walshe, Zhuode Liu, Lin Guan 等AAAI 2020 · 被引用 77 次
- Comparison of Visual Saliency for Dynamic Point Clouds: Task-free vs. Task-dependentXuemei Zhou, Irene Viola, Silvia Rossi, Pablo CésarIEEE VR 2025 · 被引用 5 次
- Predicting Human Scanpaths in Visual Question AnsweringXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2021
