Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor Data
Wei Cui, Lukai Fan, Zhenghua Chen, Min Wu, Shili Xiang, Haixia Wang, Bing Li
摘要
Sensory Temporal Action Detection (STAD) aims to localize and classify human actions within long, untrimmed sequences captured by non-visual sensors such as WiFi or inertial measurement units (IMUs). Unlike video-based TAD, STAD poses unique challenges due to the low-dimensional, noisy, and heterogeneous nature of sensory data, as well as the real-time and resource constraints on edge devices. While recent STAD models have improved detection performance, their high computational cost hampers practical deployment. In this paper, we propose SlimSTAD, a simple yet effective framework that achieves both high accuracy and low latency for STAD. SlimSTAD features a novel Decoupled Channel Modeling (DCM) encoder, which preserves modality-specific temporal features and enables efficient inter-channel aggregation via lightweight graph attention. An anchor-free cascade predictor then refines action boundaries and class predictions in a two-stage design without dense proposals. Experiments on two real-world datasets demonstrate that SlimSTAD outperforms strong video-derived and sensory baselines by an average of 2.1 mAP, while significantly reducing GFLOPs, parameters, and latency, validating its effectiveness for real-world, edge-aware STAD deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- WiFi CSI Based Temporal Activity Detection via Dual Pyramid NetworkZhendong Liu, Le Zhang, Bing Li, Yingjie Zhou 等AAAI 2025 · 被引用 6 次
- XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and GlassesBo Lan, Pei Li, Jiaxi Yin, Yunpeng Song 等UbiComp 2025 · 被引用 6 次
- Learning Salient Boundary Feature for Anchor-free Temporal Action LocalizationChuming Lin, Chengming Xu, Donghao Luo, Yabiao Wang 等CVPR 2021
相关 Paper
- MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity RecognitionZiqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing 等UbiComp 2023 · 被引用 22 次
- SEGALL: A Unified Active Learning Framework for Wireless Sensing Data SegmentationNaiyu Zheng, Ruofeng Liu, Xiaoyi Fan, Cong Zhang 等UbiComp 2025 · 被引用 3 次
- TriDet: Temporal Action Detection with Relative Boundary ModelingDingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma 等CVPR 2023
- Text-Infused Attention and Foreground-Aware Modeling for Zero-Shot Temporal Action DetectionYearang Lee, Ho-Joong Kim, Seong-Whan LeeNeurIPS 2024 · 被引用 12 次
- Listen to Look: Action Recognition by Previewing AudioRuohan Gao, Tae-Hyun Oh, Kristen Grauman, Lorenzo TorresaniCVPR 2020
