MEDIRL: Predicting the Visual Attention of Drivers via Maximum Entropy Deep Inverse Reinforcement Learning
Sonia Baee, Erfan Pakdamanian, Inki Kim, Lu Feng, Vicente Ordonez, Laura E. Barnes
Abstract
Inspired by human visual attention, we propose a novel inverse reinforcement learning formulation using Maximum Entropy Deep Inverse Reinforcement Learning (MEDIRL) for predicting the visual attention of drivers in accident-prone situations. MEDIRL predicts fixation locations that lead to maximal rewards by learning a task-sensitive reward function from eye fixation patterns recorded from attentive drivers. Additionally, we introduce EyeCar, a new driver attention dataset in accident-prone situations. We conduct comprehensive experiments to evaluate our proposed model on three common benchmarks: (DR(eye)VE, BDD-A, DADA-2000), and our EyeCar dataset. Results indicate that MEDIRL outperforms existing models for predicting attention and achieves state-of-the-art performance. We present extensive ablation studies to provide more insights into different features of our proposed model.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06e6fc39-96d2-456e-b354-9cd8efcc1648Cited by top-tier papers10
- Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge EmbeddingPengfei Zhu, Mengshi Qi, Xia Li, Weijian Li et al.ICCV 2023 · 23 citations
- FBLNet: FeedBack Loop Network for Driver Attention PredictionYilong Chen, Zhixiong Nan, Tao XiangICCV 2023 · 21 citations
- SalM²: An Extremely Lightweight Saliency Mamba Model for Real-Time Cognitive Awareness of Driver AttentionChunyu Zhao, Wentao Mu, Xian Zhou, Wenbo Liu et al.AAAI 2025 · 16 citations
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin et al.ICCV 2025 · 8 citations
- Cross-Modality Graph-based Language and Sensor Data Co-Learning of Human-Mobility InteractionMahan Tabatabaie, Suining He, Kang G. ShinUbiComp 2023 · 6 citations
Builds on11
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 666 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 189 citations
- DGaze: CNN-Based Gaze Prediction in Dynamic ScenesZhiming Hu, Sheng Li, Congyi Zhang, Kangrui Yi et al.IEEE VR 2020 · 105 citations
Related papers
- DRIVE: Deep Reinforced Accident Anticipation with Visual ExplanationWentao Bao, Qi Yu, Yu KongICCV 2021 · 64 citations
- Predicting Goal-Directed Human Attention Using Inverse Reinforcement LearningZhibo Yang, Lihan Huang, Yupei Chen, Zijun Wei et al.CVPR 2020
- GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory PredictionMuleilan Pei, Shaoshuai Shi, Lu Zhang, Peiliang Li et al.ICML 2025
- From Gaze to Movement: Predicting Visual Attention for Autonomous Driving Human-Machine Interaction based on Programmatic Imitation LearningYexin Huang, Yongbin Lin, Lishengsa Yue, Zhihong Yao et al.ICCV 2025 · 2 citations
- PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point PredictionHanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao et al.ACM MM 2025
