From Gaze to Movement: Predicting Visual Attention for Autonomous Driving Human-Machine Interaction based on Programmatic Imitation Learning
Yexin Huang, Yongbin Lin, Lishengsa Yue, Zhihong Yao, Jie Wang
摘要
Human-machine interaction technology requires not only the distribution of human visual attention but also the prediction of the gaze point trajectory. We introduce PILOT, a programmatic imitation learning approach that predicts a driver's eye movements based on a set of rule-based conditions. These conditions-derived from driving operations and traffic flow characteristics-define how gaze shifts occur. They are initially identified through incremental synthesis, a heuristic search method, and then refined via L-BFGS, a numerical optimization technique. These humanreadable rules enable us to understand drivers' eye movement patterns and make efficient and explainable predictions. We also propose DATAD, a dataset that covers 12 types of autonomous driving takeover scenarios, collected from 60 participants and comprising approximately 600,000 frames of gaze point data. Compared to existing eye-tracking datasets, DATAD includes additional driving metrics and surrounding traffic flow characteristics, providing richer contextual information for modeling gaze behavior. Experimental evaluations of PILOT on DATAD demonstrate superior accuracy and faster prediction speeds compared to four baseline models. Specifically, PILOT reduces the MSE of predicted trajectories by 38.59% to 88.02% and improves the accuracy of gaze object predictions by 6.90% to 55.06%. Moreover, PILOT achieves these gains with approximately 30% lower prediction time, offering both more accurate and more efficient eye movement prediction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationXiaoyu Shi, Zhaoyang Huang, Weikang Bian, Dasong Li 等ICCV 2023 · 被引用 112 次
- Pyramid Grafting Network for One-Stage High Resolution Saliency DetectionChenxi Xie, Changqun Xia, Mingcan Ma, Zhirui Zhao 等CVPR 2022 · 被引用 112 次
- MEDIRL: Predicting the Visual Attention of Drivers via Maximum Entropy Deep Inverse Reinforcement LearningSonia Baee, Erfan Pakdamanian, Inki Kim, Lu Feng 等ICCV 2021 · 被引用 65 次
- DRIVE: Deep Reinforced Accident Anticipation with Visual ExplanationWentao Bao, Qi Yu, Yu KongICCV 2021 · 被引用 64 次
- FBLNet: FeedBack Loop Network for Driver Attention PredictionYilong Chen, Zhixiong Nan, Tao XiangICCV 2023 · 被引用 21 次
相关 Paper
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin 等ICCV 2025 · 被引用 8 次
- PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory PredictionAmir Rasouli, Iuliia Kotseruba, Toni Kunic, John K. TsotsosICCV 2019 · 被引用 411 次
- DriverGaze360: OmniDirectional Driver Attention with Object-Level GuidanceShreedhar Govil, Didier Stricker, Jason R. RambachCVPR 2026
- Leveraging Driver Field-of-View for Multimodal Ego-Trajectory PredictionM. Eren Akbiyik, Nedko Savov, Danda Pani Paudel, Nikola Popovic 等ICLR 2025
- Atari-HEAD: Atari Human Eye-Tracking and Demonstration DatasetRuohan Zhang, Calen Walshe, Zhuode Liu, Lin Guan 等AAAI 2020 · 被引用 77 次
