Lightweight Spatio-Temporal Modeling via Temporally Shifted Distillation for Real-Time Accident Anticipation
Patrik Patera, Yie-Tarng Chen, Wen-Hsien Fang
Abstract
Anticipating traffic accidents in real time is critical for intelligent transportation systems, yet remains challenging under edge-device constraints. We propose a lightweight spatio-temporal framework that introduces a temporally shifted distillation strategy, enabling a student model to acquire predictive temporal dynamics from a frozen image-based teacher without requiring a video pre-trained teacher. The student combines a RepMixer spatial encoding with a RWKV-inspired recurrent module for efficient long-range temporal reasoning. To enhance robustness under partial observability, we design a masking memory strategy that leverages memory retention to reconstruct missing visual tokens, effectively simulating occlusions and future events. In addition, multi-modal vision-language supervision enriches semantic grounding. Our framework achieves state-of-the-art performance on multiple real-world dashcam benchmarks while sustaining real-time inference on resource-limited platforms such as the NVIDIA Jetson Orin Nano. Remarkably, it is 3-7 smaller than leading approaches yet delivers superior accuracy and earlier anticipation, underscoring its practicality for deployment in intelligent vehicles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1dae9b0-a5c5-4bfe-9aa6-818095ba7cffBuilds on13
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly DetectionPeng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou et al.AAAI 2024 · 220 citations
Related papers
- Ultrafast Video Attention Prediction with Coupled Knowledge DistillationKui Fu, Peipei Shi, Yafei Song, Shiming Ge et al.AAAI 2020 · 11 citations
- AMap: Distilling Future Priors for Ahead-Aware Online HD Map ConstructionRuikai Li, Xinrun Li, Mengwei Xie, Hao Shan et al.CVPR 2026 · 8 citations
- Online Model Distillation for Efficient Video InferenceRavi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan et al.ICCV 2019 · 131 citations
- Learning Lightweight Object Detectors via Multi-Teacher Progressive DistillationShengcao Cao, Mengtian Li, James Hays, Deva Ramanan et al.ICML 2023 · 17 citations
- How many Observations are Enough? Knowledge Distillation for Trajectory ForecastingAlessio Monti, Angelo Porrello, Simone Calderara, Pasquale Coscia et al.CVPR 2022 · 62 citations
