Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous Vehicles
Manuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne, Simon Reiß, Michael Voit, Rainer Stiefelhagen
Abstract
We introduce the novel domain-specific Drive&Act benchmark for fine-grained categorization of driver behavior. Our dataset features twelve hours and over 9.6 million frames of people engaged in distractive activities during both, manual and automated driving. We capture color, infrared, depth and 3D body pose information from six views and densely label the videos with a hierarchical annotation scheme, resulting in 83 categories. The key challenges of our dataset are: (1) recognition of fine-grained behavior inside the vehicle cabin; (2) multi-modal activity recognition, focusing on diverse data streams; and (3) a cross view recognition benchmark, where a model handles data from an unfamiliar domain, as sensor type and placement in the cabin can change between vehicles. Finally, we provide challenging benchmarks by adopting prominent methods for video- and body pose-based action recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47f36e77-f5d4-4deb-8695-2f55f75f7eb4Cited by top-tier papers11
- AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionDingkang Yang, Shuai Huang, Zhi Xu, Zhenpeng Li et al.ICCV 2023 · 72 citations
- AutoVis: Enabling Mixed-Immersive Analysis of Automotive User Interface Interaction StudiesPascal Jansen, Julian Britten, Alexander Häusele, Thilo Segschneider et al.CHI 2023 · 38 citations
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMsYumin Choi, Dongki Kim, Jinheon Baek, Sung Ju HwangICLR 2026 · 4 citations
- AutoTherm: A Dataset and Benchmark for Thermal Comfort Estimation Indoors and in VehiclesMark Colley, Sebastian Hartwig, Albin Zeqiri, Timo Ropinski et al.UbiComp 2024 · 4 citations
- Frame2Freq: Spectral Adapters for Fine-Grained Video UnderstandingThinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina RoitbergCVPR 2026 · 2 citations
Related papers
- UAV-Human: A Large Benchmark for Human Behavior Understanding With Unmanned Aerial VehiclesTianjiao Li, Jun Liu, Wei Zhang, Yun Ni et al.CVPR 2021
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He et al.CVPR 2022 · 168 citations
- Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous DrivingAlexey Nekrasov, Malcolm Burdorf, Stewart Worrall, Bastian Leibe et al.CVPR 2025
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt et al.ICCV 2019 · 108 citations
- DriverGaze360: OmniDirectional Driver Attention with Object-Level GuidanceShreedhar Govil, Didier Stricker, Jason R. RambachCVPR 2026
