MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity Recognition
Ziqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing, Shwetak N. Patel, Xin Liu, Yuanchun Shi
摘要
Multimodal sensors provide complementary information to develop accurate machine-learning methods for human activity recognition (HAR), but introduce significantly higher computational load, which reduces efficiency. This paper proposes an efficient multimodal neural architecture for HAR using an RGB camera and inertial measurement units (IMUs) called Multimodal Temporal Segment Attention Network (MMTSA). MMTSA first transforms IMU sensor data into a temporal and structure-preserving gray-scale image using the Gramian Angular Field (GAF), representing the inherent properties of human activities. MMTSA then applies a multimodal sparse sampling method to reduce data redundancy. Lastly, MMTSA adopts an inter-segment attention module for efficient multimodal fusion. Using three well-established public datasets, we evaluated MMTSA's effectiveness and efficiency in HAR. Results show that our method achieves superior performance improvements (11.13% of cross-subject F1-score on the MMAct dataset) than the previous state-of-the-art (SOTA) methods. The ablation study and analysis suggest that MMTSA's effectiveness in fusing multimodal data for accurate HAR. The efficiency evaluation on an edge device showed that MMTSA achieved significantly better accuracy, lower computational load, and lower inference latency than SOTA methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 被引用 50 次
- Past, Present, and Future of Sensor-based Human Activity Recognition Using Wearables: A Surveying Tutorial on a Still Challenging TaskHarish Haresamudram, Chi Ian Tang, Sungho Suh, Paul Lukowicz 等UbiComp 2025 · 被引用 31 次
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao 等UbiComp 2024 · 被引用 23 次
- Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution ImagesYuntao Wang, Zirui Cheng, Xin Yi, Yan Kong 等CHI 2023 · 被引用 10 次
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew 等UbiComp 2026 · 被引用 1 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
相关 Paper
- IF-ConvTransformer: A Framework for Human Activity Recognition Using IMU Fusion and ConvTransformerYe Zhang, Longguang Wang, Huiling Chen, Aosheng Tian 等UbiComp 2022 · 被引用 47 次
- Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor DataWei Cui, Lukai Fan, Zhenghua Chen, Min Wu 等AAAI 2026
- GIobalFusion: A Global Attentional Deep Learning Framework for Multisensor Information FusionShengzhong Liu, Shuochao Yao, Jinyang Li, Dongxin Liu 等UbiComp 2020 · 被引用 57 次
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal AttentionKatsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro OkadaACM MM 2021 · 被引用 9 次
- Cosmo: contrastive fusion learning with small data for multimodal human activity recognitionXiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi 等MobiCom 2022 · 被引用 94 次
