Accurate and Steady Inertial Pose Estimation through Sequence Structure Learning and Modulation
Yinghao Wu, Chaoran Wang, Lu Yin, Shihui Guo, Yipeng Qin
Abstract
Transformer models excel at capturing long-range dependencies in sequential data, but lack explicit mechanisms to leverage structural patterns inherent in fixed-length input sequences. In this paper, we propose a novel sequence structure learning and modulation approach that endows Transformers with the ability to model and utilize such fixed-sequence structural properties for improved performance on inertial pose estimation tasks. Specifically, our method introduces a Sequence Structure Module (SSM) that utilizes structural information of fixed-length inertial sensor readings to adjust the input features of transformers. Such structural information can either be acquired by learning or specified based on users’ prior knowledge. To justify the prospect of our approach, we show that i) injecting spatial structural information of IMUs/joints learned from data improves accuracy, while ii) injecting temporal structural information based on smooth priors reduces jitter (i.e., improves steadiness), in a spatial-temporal transformer solution for inertial pose estimation. Extensive experiments across multiple benchmark datasets demonstrate the superiority of our approach against state-of-the-art methods and has the potential to advance the design of the transformer architecture for fixed-length sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aeb73899-9510-4c29-9ebc-56def80798c2Cited by top-tier papers7
- ToF-IP: Time-of-Flight Enhanced Sparse Inertial Poser for Real-time Human Motion CaptureYuan Yao, Shifan Jiang, Yangqing Hou, Chengxu Zuo et al.NeurIPS 2025 · 2 citations
- MagShield: Towards Better Robustness in Sparse Inertial Motion Capture Under Magnetic DisturbancesYunzhe Shao, Xinyu Yi, Lu Yin, Shihui Guo et al.ICCV 2025 · 1 citation
- MODA: Motion-Drift Augmentation for Inertial Human Motion AnalysisYinghao Wu, Shihui Guo, Yipeng QinCVPR 2025
- Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor DistancesDominik Hollidt, Tommaso Bendinelli, Christian HolzCVPR 2026
- Improving Sparse IMU-based Motion Capture with Motion Label SmoothingZhaorui Meng, Lu Yin, Yangqing Hou, Anjun Chen et al.AAAI 2026
Builds on28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series ForecastingHaixu Wu, Jiehui Xu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 5,824 citations
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang et al.ICML 2022 · 2,912 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- Sequential Joint Dependency Aware Human Pose Estimation with State Space ModelHanxi Yin, Shaodi You, Jungong Han, Zhixiang ChenAAAI 2025 · 1 citation
- iMoT: Inertial Motion Transformer for Inertial NavigationSon Minh Nguyen, Duc Viet Le, Paul J. M. HavingaAAAI 2025 · 2 citations
- CTIN: Robust Contextual Transformer Network for Inertial NavigationBingbing Rao, Ehsan Kazemi, Yifan Ding, Devu M. Shila et al.AAAI 2022 · 66 citations
- HiPoser: 3D Human Pose Estimation with Hierarchical Shared Learning at Parts-Level Using Inertial Measurement UnitsGuorui Liao, Chunyuan Zheng, Li Cheng, Haoyu Xie et al.AAAI 2025
- Spatial-Related Sensors Matters: 3D Human Motion Reconstruction Assisted with Textual SemanticsXueyuan Yang, Chao Yao, Xiaojuan BanAAAI 2024 · 4 citations
