DualScope: Capturing Critical Spatial and Temporal Cues for Distracted Driving Activity Recognition
Zhijie Qiu, Shuaibo Li, Laixin Zhang, Xuming Hu, Wei Ma
摘要
Accurately recognizing distracted driving activities in realworld scenarios is essential for improving road and pedestrian safety. However, existing approaches are prone to attend to irrelevant scene contexts and are susceptible to interference from redundant frames, compromising their robustness in complex driving environments. To overcome these limitations, we propose DualScope, a novel framework that captures behaviorally critical information from spatial and temporal perspectives. In the Spatial Scope, we introduce a Synergistic Behavior-Centric Distillation mechanism that leverages two key information sources: (1) position-aware knowledge derived from the SAM model that enhances the perception of critical regions and their semantic interaction structures and (2) fine-grained visual details obtained from cropped key regions that improve the model's ability to capture detailed patterns in behavior-relevant areas. In the Temporal Scope, we present the Saliency-Aware Fine-to-Coarse Temporal Modeling module comprising three components: a Fine-Grained Motion Encoder for capturing local inter-frame dependencies, a Dynamic Difference Extractor for extracting salient motion dynamics, and a Saliency-Aware Temporal Pyramid Mamba for integrating these features to enable multiscale temporal modeling. This design effectively captures short-term motions and long-term behavioral patterns. Furthermore, incorporating salient dynamics enhances the model's focus on substantial behavioral variations. Extensive experiments on seven publicly available distracted driving activity recognition datasets demonstrate that DualScope consistently outperforms state-of-the-art methods, validating its effectiveness in capturing behavioral cues across spatial and temporal dimensions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
相关 Paper
- Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesManuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne 等ICCV 2019 · 被引用 235 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
- Temporal Pyramid Network for Pedestrian Trajectory Prediction with Multi-SupervisionRongqin Liang, Yuanman Li, Xia Li, Yi Tang 等AAAI 2021 · 被引用 55 次
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
- FIFA: Fine-grained Inter-frame Attention for Driver's Video Gaze EstimationDaosong Hu, Mingyue Cui, Kai HuangCVPR 2025
