DualScope: Capturing Critical Spatial and Temporal Cues for Distracted Driving Activity Recognition
Zhijie Qiu, Shuaibo Li, Laixin Zhang, Xuming Hu, Wei Ma
Abstract
Accurately recognizing distracted driving activities in realworld scenarios is essential for improving road and pedestrian safety. However, existing approaches are prone to attend to irrelevant scene contexts and are susceptible to interference from redundant frames, compromising their robustness in complex driving environments. To overcome these limitations, we propose DualScope, a novel framework that captures behaviorally critical information from spatial and temporal perspectives. In the Spatial Scope, we introduce a Synergistic Behavior-Centric Distillation mechanism that leverages two key information sources: (1) position-aware knowledge derived from the SAM model that enhances the perception of critical regions and their semantic interaction structures and (2) fine-grained visual details obtained from cropped key regions that improve the model's ability to capture detailed patterns in behavior-relevant areas. In the Temporal Scope, we present the Saliency-Aware Fine-to-Coarse Temporal Modeling module comprising three components: a Fine-Grained Motion Encoder for capturing local inter-frame dependencies, a Dynamic Difference Extractor for extracting salient motion dynamics, and a Saliency-Aware Temporal Pyramid Mamba for integrating these features to enable multiscale temporal modeling. This design effectively captures short-term motions and long-term behavioral patterns. Furthermore, incorporating salient dynamics enhances the model's focus on substantial behavioral variations. Extensive experiments on seven publicly available distracted driving activity recognition datasets demonstrate that DualScope consistently outperforms state-of-the-art methods, validating its effectiveness in capturing behavioral cues across spatial and temporal dimensions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef745e1b-f7d7-44c4-b32a-89e0caa8a827Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
Related papers
- Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesManuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne et al.ICCV 2019 · 235 citations
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 217 citations
- Temporal Pyramid Network for Pedestrian Trajectory Prediction with Multi-SupervisionRongqin Liang, Yuanman Li, Xia Li, Yi Tang et al.AAAI 2021 · 55 citations
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
- FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze EstimationDaosong Hu, Mingyue Cui, Kai HuangCVPR 2025
