X-Fi: A Modality-Invariant Foundation Model for Multimodal Human Sensing
Xinyan Chen, Jianfei Yang
Abstract
Human sensing, which employs various sensors and advanced deep learning technologies to accurately capture and interpret human body information, has significantly impacted fields like public security and robotics. However, current human sensing primarily depends on modalities such as cameras and LiDAR, each of which has its own strengths and limitations. Furthermore, existing multimodal fusion solutions are typically designed for fixed modality combinations, requiring extensive retraining when modalities are added or removed for diverse scenarios. In this paper, we propose a modality-invariant foundation model for all modalities, X-Fi, to address these issues. X-Fi enables the independent or combinatory use of sensor modalities without additional training by utilizing a transformer structure to accommodate variable input sizes and incorporating a novel "X-fusion" mechanism to preserve modality-specific features during multimodal integration. This approach not only enhances adaptability but also facilitates the learning of complementary features across modalities. Extensive experiments conducted on the MM-Fi and XRF55 datasets, employing six distinct modalities, demonstrate that X-Fi achieves state-of-the-art performance in human pose estimation (HPE) and human activity recognition (HAR) tasks. The findings indicate that our proposed model can efficiently support a wide range of human sensing applications, ultimately contributing to the evolution of scalable, multimodal sensing technologies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4876871-3d72-418a-aab9-99f6ba89e271Cited by top-tier papers5
- Towards Balanced Multi-Modal Learning in 3D Human Pose EstimationMengshi Qi, Jiaxuan Peng, Xianlin Zhang, Huadong MaCVPR 2026 · 12 citations
- HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and ReasoningChuhao Zhou, Jianfei YangNeurIPS 2025 · 4 citations
- BriMA: Bridged Modality Adaptation for Multi-Modal Continual Action Quality AssessmentKanglei Zhou, Chang Li, Qingyi Pan, Liyuan WangCVPR 2026 · 3 citations
- RF-MatID: Dataset and Benchmark for Radio Frequency Material IdentificationXinyan Chen, Qinchun Li, Ruiqin Ma, Jiaqi Bai et al.ICLR 2026 · 1 citation
- XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the EdgeYu Zhang, Xi Zhang, Hualin zhou, Xinyuan Chen et al.ICML 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- Are Multimodal Transformers Robust to Missing Modality?Mengmeng Ma, Jian Ren, Long Zhao, Davide Testuggine et al.CVPR 2022 · 153 citations
Related papers
- MASTER: A Multi-modal Foundation Model for Human Activity RecognitionGuanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao et al.UbiComp 2025 · 8 citations
- Weakly Supervised 3D Multi-Person Pose Estimation for Large-Scale Scenes Based on Monocular Camera and Single LiDARPeishan Cong, Yiteng Xu, Yiming Ren, Juze Zhang et al.AAAI 2023 · 37 citations
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu et al.ICML 2024 · 64 citations
- FreeCap: Hybrid Calibration-Free Motion Capture in Open EnvironmentsAoru Xue, Yiming Ren, Zining Song, Mao Ye et al.AAAI 2025 · 4 citations
- SATPose: Improving Monocular 3D Pose Estimation with Spatial-aware Ground TactilityLishuang Zhan, Enting Ying, Jiabao Gan, Shihui Guo et al.ACM MM 2024 · 2 citations
