Vistar: Enhancing the Perception Capability of LLMs under Imprecise IMU-Text Alignment
Yatong Chen, Chenzhi Hu, Bowen He, Ruijie Wang, Xiaomin Ouyang, Shengzhong Liu, Jianxin Li, Fan Wu, Guihai Chen
Abstract
This paper introduces Vistar, a novel self-supervised framework for inertial measurement unit (IMU) signal perception designed for large language models (LLMs). Unlike visual data, IMU signals are high-frequency time series with low interpretability, making manual annotation with natural language particularly challenging. Even when using vision-language models (VLMs) to describe events in videos synchronized with IMU signals, a semantic gap remains between high-level visual semantics and low-level IMU vibrations. The core idea of Vistar is to achieve accurate IMU signal perception through collaborations between offline cross-modal alignment and online retrieval-augmented generation. During offline training, Vistar uses pretrained vision and language encoders as anchors to learn IMU encoders via hierarchical cross-modal contrastive learning, establishing both inter- and intra-sample alignment. Given that the enhanced training strategy still fails to achieve precise alignment between IMU and text, during online inference, Vistar further employs a retrieval-augmented generation mechanism to generate distilled textual descriptions from similar text filtered based on structural relations of their paired IMU samples. Extensive evaluations on three multimodal datasets demonstrate that Vistar consistently outperforms state-of-the-art (SOTA) baselines by up to 57.45% in IMU-to-text retrieval and improves the generated text similarity with ground truths in IMU perception by up to 31.90%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a02bbb3a-6dce-4737-955c-3c4c914110e8Related papers
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 1 citation
- IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity RecognitionZikang Leng, Amitrajit Bhattacharjee, Hrudhai Rajasekhar, Lizhe Zhang et al.UbiComp 2024 · 59 citations
- COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-trainingSanghwan Kim, Rui Xiao, Mariana-Iuliana Georgescu, Stephan Alaniz et al.CVPR 2025
- Head2Body: Body Pose Generation from Multi-Sensory Head-Mounted InputsMinh Tran, Hongda Mao, Qingshuang Chen, Yelin KimICCV 2025 · 1 citation
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
