IF-ConvTransformer: A Framework for Human Activity Recognition Using IMU Fusion and ConvTransformer
Ye Zhang, Longguang Wang, Huiling Chen, Aosheng Tian, Shilin Zhou, Yulan Guo
Abstract
Recent advances in sensor based human activity recognition (HAR) have exploited deep hybrid networks to improve the performance. These hybrid models combine Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) to leverage their complementary advantages, and achieve impressive results. However, the roles and associations of different sensors in HAR are not fully considered by these models, leading to insufficient multi-modal fusion. Besides, the commonly used RNNs in HAR suffer from the 'forgetting' defect, which raises difficulties in capturing long-term information. To tackle these problems, an HAR framework composed of an Inertial Measurement Unit (IMU) fusion block and an applied ConvTransformer subnet is proposed in this paper. Inspired by the complementary filter, our IMU fusion block performs multi-modal fusion of commonly used sensors according to their physical relationships. Consequently, the features of different modalities can be aggregated more effectively. Then, the extracted features are fed into the applied ConvTransformer subnet for classification. Thanks to its convolutional subnet and self-attention layers, ConvTransformer can better capture local features and construct long-term dependencies. Extensive experiments on eight benchmark datasets demonstrate the superior performance of our framework. The source code will be published soon.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers8
- AutoAugHAR: Automated Data Augmentation for Sensor-based Human Activity RecognitionYexu Zhou, Haibin Zhao, Yiran Huang, Tobias Röddiger et al.UbiComp 2024 · 25 citations
- Temporal Action Localization for Inertial-based Human Activity RecognitionMarius Bock, Michael Möller, Kristof Van LaerhovenUbiComp 2025 · 14 citations
- CALANet: Cheap All-Layer Aggregation for Human Activity RecognitionJaegyun Park, Dae-Won Kim, Jaesung LeeNeurIPS 2024 · 11 citations
- Deep Heterogeneous Contrastive Hyper-Graph Learning for In-the-Wild Context-Aware Human Activity RecognitionWen Ge, Guanyi Mou, Emmanuel O. Agu, Kyumin LeeUbiComp 2024 · 9 citations
- IMUCoCo: Enabling Flexible On-Body IMU Placement for Human Pose Estimation and Activity RecognitionHaozhe Zhou, Riku Arakawa, Yuvraj Agarwal, Mayank GoelUIST 2025 · 5 citations
Related papers
- MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity RecognitionZiqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing et al.UbiComp 2023 · 22 citations
- rTsfNet: A DNN Model with Multi-head 3D Rotation and Time Series Feature Extraction for IMU-based Human Activity RecognitionYu EnokiboriUbiComp 2025 · 10 citations
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionHao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang et al.NeurIPS 2025 · 11 citations
- GIobalFusion: A Global Attentional Deep Learning Framework for Multisensor Information FusionShengzhong Liu, Shuochao Yao, Jinyang Li, Dongxin Liu et al.UbiComp 2020 · 57 citations
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew et al.UbiComp 2026 · 1 citation
