Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous Vehicles
Manuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne, Simon Reiß, Michael Voit, Rainer Stiefelhagen
摘要
We introduce the novel domain-specific Drive&Act benchmark for fine-grained categorization of driver behavior. Our dataset features twelve hours and over 9.6 million frames of people engaged in distractive activities during both, manual and automated driving. We capture color, infrared, depth and 3D body pose information from six views and densely label the videos with a hierarchical annotation scheme, resulting in 83 categories. The key challenges of our dataset are: (1) recognition of fine-grained behavior inside the vehicle cabin; (2) multi-modal activity recognition, focusing on diverse data streams; and (3) a cross view recognition benchmark, where a model handles data from an unfamiliar domain, as sensor type and placement in the cabin can change between vehicles. Finally, we provide challenging benchmarks by adopting prominent methods for video- and body pose-based action recognition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionDingkang Yang, Shuai Huang, Zhi Xu, Zhenpeng Li 等ICCV 2023 · 被引用 72 次
- AutoVis: Enabling Mixed-Immersive Analysis of Automotive User Interface Interaction StudiesPascal Jansen, Julian Britten, Alexander Häusele, Thilo Segschneider 等CHI 2023 · 被引用 38 次
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMsYumin Choi, Dongki Kim, Jinheon Baek, Sung Ju HwangICLR 2026 · 被引用 4 次
- AutoTherm: A Dataset and Benchmark for Thermal Comfort Estimation Indoors and in VehiclesMark Colley, Sebastian Hartwig, Albin Zeqiri, Timo Ropinski 等UbiComp 2024 · 被引用 4 次
- Frame2Freq: Spectral Adapters for Fine-Grained Video UnderstandingThinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina RoitbergCVPR 2026 · 被引用 2 次
相关 Paper
- UAV-Human: A Large Benchmark for Human Behavior Understanding With Unmanned Aerial VehiclesTianjiao Li, Jun Liu, Wei Zhang, Yun Ni 等CVPR 2021
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He 等CVPR 2022 · 被引用 168 次
- Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous DrivingAlexey Nekrasov, Malcolm Burdorf, Stewart Worrall, Bastian Leibe 等CVPR 2025
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt 等ICCV 2019 · 被引用 108 次
- DriverGaze360: OmniDirectional Driver Attention with Object-Level GuidanceShreedhar Govil, Didier Stricker, Jason R. RambachCVPR 2026
