Learning from Observer Gaze: Zero-Shot Attention Prediction Oriented by Human-Object Interaction Recognition
Yuchen Zhou, Linkai Liu, Chao Gou
摘要
Most existing attention prediction research focuses on salient instances like humans and objects. However, the more complex interaction-oriented attention, arising from the comprehension of interactions between instances by human observers, remains largely unexplored. This is equally crucial for advancing human-machine interaction and human-centered artificial intelligence. To bridge this gap, we first collect a novel gaze fixation dataset named IG, comprising 530,000 fixation points across 740 diverse in-teraction categories, capturing visual attention during human observers' cognitive processes of interactions. Subsequently, we introduce the zero-shot interaction-oriented attention prediction task (ZeroIA), which challenges models to predict visual cues for interactions not encountered during training. Thirdly, we present the Interactive Attention model (IA), designed to emulate human observers' cognitive processes to tackle the ZeroIA problem. Extensive experiments demonstrate that the proposed IA outperforms other state-of-the-art approaches in both ZeroIA and fully supervised settings. Lastly, we endeavor to apply interaction-oriented attention to the interaction recognition task itself. Further experimental results demonstrate the promising potential to enhance the performance and inter-pretability of existing state-of-the-art HOI models by incorporating real human attention data from IG and attention labels generated by IA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG DecodingYuchen Zhou, Jiamin Wu, Zichen Ren, Zhouheng Yao 等NeurIPS 2025 · 被引用 71 次
- Gaze-VLM: Bridging Gaze and VLMs through Attention Regularization for Egocentric UnderstandingAnupam Pani, Yanchao YangNeurIPS 2025 · 被引用 12 次
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin 等ICCV 2025 · 被引用 8 次
- Gaze Beyond the Frame: Forecasting Egocentric 3D Visual SpanHeeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock 等NeurIPS 2025 · 被引用 3 次
- Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsYuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Spatially Conditioned Graphs for Detecting Human-Object InteractionsFrederic Z. Zhang, Dylan Campbell, Stephen GouldICCV 2021 · 被引用 170 次
- The Role of Eye Gaze in Security and Privacy Applications: Survey and Future HCI Research DirectionsChristina P. Katsini, Yasmeen Abdrabou, George E. Raptis, Mohamed Khamis 等CHI 2020 · 被引用 167 次
- GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionYue Liao, Aixi Zhang, Miao Lu, Yongliang Wang 等CVPR 2022 · 被引用 136 次
相关 Paper
- Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human AttentionSounak Mondal, Zhibo Yang, Seoyoung Ahn, Dimitris Samaras 等CVPR 2023
- Detecting Human-Object Interactions via Functional GeneralizationAnkan Bansal, Sai Saketh Rambhatla, Abhinav Shrivastava, Rama ChellappaAAAI 2020 · 被引用 131 次
- Goal-Oriented Gaze Estimation for Zero-Shot LearningYang Liu, Lei Zhou, Xiao Bai, Yifei Huang 等CVPR 2021
- Learning Human-Object Interaction Detection Using Interaction PointsTiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan 等CVPR 2020
- Reformulating HOI Detection As Adaptive Set PredictionMingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen 等CVPR 2021
