MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity Recognition
Ziqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing, Shwetak N. Patel, Xin Liu, Yuanchun Shi
Abstract
Multimodal sensors provide complementary information to develop accurate machine-learning methods for human activity recognition (HAR), but introduce significantly higher computational load, which reduces efficiency. This paper proposes an efficient multimodal neural architecture for HAR using an RGB camera and inertial measurement units (IMUs) called Multimodal Temporal Segment Attention Network (MMTSA). MMTSA first transforms IMU sensor data into a temporal and structure-preserving gray-scale image using the Gramian Angular Field (GAF), representing the inherent properties of human activities. MMTSA then applies a multimodal sparse sampling method to reduce data redundancy. Lastly, MMTSA adopts an inter-segment attention module for efficient multimodal fusion. Using three well-established public datasets, we evaluated MMTSA's effectiveness and efficiency in HAR. Results show that our method achieves superior performance improvements (11.13% of cross-subject F1-score on the MMAct dataset) than the previous state-of-the-art (SOTA) methods. The ablation study and analysis suggest that MMTSA's effectiveness in fusing multimodal data for accurate HAR. The efficiency evaluation on an edge device showed that MMTSA achieved significantly better accuracy, lower computational load, and lower inference latency than SOTA methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da375d20-bcd5-46ab-ae8f-fa589018f2d7Cited by top-tier papers6
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 50 citations
- Past, Present, and Future of Sensor-based Human Activity Recognition Using Wearables: A Surveying Tutorial on a Still Challenging TaskHarish Haresamudram, Chi Ian Tang, Sungho Suh, Paul Lukowicz et al.UbiComp 2025 · 31 citations
- G-VOILA: Gaze-Facilitated Information Querying in Daily ScenariosZeyu Wang, Yuanchun Shi, Yuntao Wang, Yuchen Yao et al.UbiComp 2024 · 23 citations
- Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution ImagesYuntao Wang, Zirui Cheng, Xin Yi, Yan Kong et al.CHI 2023 · 10 citations
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew et al.UbiComp 2026 · 1 citation
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
Related papers
- IF-ConvTransformer: A Framework for Human Activity Recognition Using IMU Fusion and ConvTransformerYe Zhang, Longguang Wang, Huiling Chen, Aosheng Tian et al.UbiComp 2022 · 47 citations
- Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor DataWei Cui, Lukai Fan, Zhenghua Chen, Min Wu et al.AAAI 2026
- GIobalFusion: A Global Attentional Deep Learning Framework for Multisensor Information FusionShengzhong Liu, Shuochao Yao, Jinyang Li, Dongxin Liu et al.UbiComp 2020 · 57 citations
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal AttentionKatsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro OkadaACM MM 2021 · 9 citations
- Cosmo: contrastive fusion learning with small data for multimodal human activity recognitionXiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi et al.MobiCom 2022 · 94 citations
