AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception
Dingkang Yang, Shuai Huang, Zhi Xu, Zhenpeng Li, Shunli Wang, Mingcheng Li, Yuzheng Wang, Yang Liu, Kun Yang, Zhaoyu Chen, Yan Wang, Jing Liu
Abstract
Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent PerceptionDingkang Yang, Kun Yang, Yuzheng Wang, Jing Liu et al.NeurIPS 2023 · 160 citations
- Learning Causality-inspired Representation Consistency for Video Anomaly DetectionYang Liu, Zhaoyang Xia, Mengyang Zhao, Donglai Wei et al.ACM MM 2023 · 48 citations
- Efficient Decision-based Black-box Patch Attacks on Video RecognitionKaixun Jiang, Zhaoyu Chen, Hao Huang, Jiafeng Wang et al.ICCV 2023 · 30 citations
- Robust Emotion Recognition in Context DebiasingDingkang Yang, Kun Yang, Mingcheng Li, Shunli Wang et al.CVPR 2024 · 25 citations
- Debiased Multimodal Understanding for Human Language SequencesZhi Xu, Dingkang Yang, Mingcheng Li, Yuzheng Wang et al.AAAI 2025 · 16 citations
Builds on11
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
- Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesManuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne et al.ICCV 2019 · 235 citations
- Content-based Unrestricted Adversarial AttackZhaoyu Chen, Bo Li, Shuang Wu, Kaixun Jiang et al.NeurIPS 2023 · 132 citations
Related papers
- DriverGaze360: OmniDirectional Driver Attention with Object-Level GuidanceShreedhar Govil, Didier Stricker, Jason R. RambachCVPR 2026
- Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection TaskXiaoqing Ye, Mao Shu, Hanyu Li, Yifeng Shi et al.CVPR 2022 · 130 citations
- RCooper: A Real-world Large-scale Dataset for Roadside Cooperative PerceptionRuiyang Hao, Siqi Fan, Yingru Dai, Zhenlin Zhang et al.CVPR 2024
- Traffic Scene Parsing Through the TSP6K DatasetPeng-Tao Jiang, Yuqi Yang, Yang Cao, Qibin Hou et al.CVPR 2024
- On-Road Object Importance Estimation: A New Dataset and A Model with Multi-Fold Top-Down GuidanceZhixiong Nan, Yilong Chen, Tianfei Zhou, Tao XiangNeurIPS 2024 · 1 citation
