DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
Shreedhar Govil, Didier Stricker, Jason R. Rambach
Abstract
Predicting driver attention is a critical problem for developing explainable autonomous driving systems and understanding driver behavior in mixed human-autonomous vehicle traffic scenarios. Although significant progress has been made through large-scale driver attention datasets and deep learning architectures, existing works are constrained by narrow frontal field-of-view and limited driving diversity. Consequently, they fail to capture the full spatial context of driving environments, especially during lane changes, turns, and interactions involving peripheral objects such as pedestrians or cyclists. In this paper, we introduce DriverGaze360, a large-scale 360 field of view driver attention dataset, containing 1 million gaze-labeled frames collected from 19 human drivers, enabling comprehensive omnidirectional modeling of driver gaze behavior. Moreover, our panoramic attention prediction approach, DriverGaze360-Net, jointly learns attention maps and attended objects by employing an auxiliary semantic segmentation head. This improves spatial awareness and attention prediction across wide panoramic inputs. Extensive experiments demonstrate that DriverGaze360-Net achieves state-of-the-art attention prediction performance on multiple metrics on panoramic driving images. Dataset and method available at https://dfki-av.github.io/drivergaze360.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- LMDrive: Closed-Loop End-to-End Driving with Large Language ModelsHao Shao, Yuxuan Hu, Letian Wang, Guanglu Song et al.CVPR 2024 · 114 citations
- MEDIRL: Predicting the Visual Attention of Drivers via Maximum Entropy Deep Inverse Reinforcement LearningSonia Baee, Erfan Pakdamanian, Inki Kim, Lu Feng et al.ICCV 2021 · 65 citations
Related papers
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin et al.ICCV 2025 · 8 citations
- Beyond Scanpaths: Graph-Based Gaze Simulation in Dynamic ScenesLuke Palmer, Petar Palasek, Hazem AbdelkawyCVPR 2026
- Leveraging Driver Field-of-View for Multimodal Ego-Trajectory PredictionM. Eren Akbiyik, Nedko Savov, Danda Pani Paudel, Nikola Popovic et al.ICLR 2025
- Capturing Omni-Range Context for Omnidirectional SegmentationKailun Yang, Jiaming Zhang, Simon Reiß, Xinxin Hu et al.CVPR 2021
- Zenseact Open Dataset: A large-scale and diverse multimodal dataset for autonomous drivingMina Alibeigi, William Ljungbergh, Adam Tonderski, Georg Hess et al.ICCV 2023 · 106 citations
