HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning
Chuhao Zhou, Jianfei Yang
Abstract
Embodied agents operating in smart homes must understand human behavior through diverse sensory inputs and communicate via natural language. While Vision-Language Models (VLMs) have enabled impressive language-grounded perception, their reliance on visual data limits robustness in real-world scenarios with occlusions, poor lighting, or privacy constraints. In this paper, we introduce HoloLLM, a Multimodal Large Language Model (MLLM) that integrates uncommon but powerful sensing modalities, such as LiDAR, infrared, mmWave radar, and WiFi, to enable seamless human perception and reasoning across heterogeneous environments. We address two key challenges: (1) the scarcity of aligned modality-text data for rare sensors, and (2) the heterogeneity of their physical signal representations. To overcome these, we design a Universal Modality-Injection Projector (UMIP) that enhances pre-aligned modality embeddings with fine-grained, text-aligned features from tailored encoders via coarse-to-fine cross-attention without introducing significant alignment overhead. We further introduce a human-VLM collaborative data curation pipeline to generate paired textual annotations for sensing datasets. Extensive experiments on two newly constructed benchmarks show that HoloLLM significantly outperforms existing MLLMs, improving language-grounded human sensing accuracy by up to 30%. This work establishes a new foundation for real-world, language-informed multisensory embodied intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29ab7614-58fa-4918-b48c-c97c3261f612Cited by top-tier papers2
- REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?Chenxi Jiang, Chuhao Zhou, Jianfei YangICLR 2026 · 8 citations
- R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal SpaceTin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric SaxCVPR 2026 · 2 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- OneLLM: One Framework to Align All Modalities with LanguageJiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang et al.CVPR 2024
- RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-wave Point Cloud SequenceZengyuan Lai, Jiarui Yang, Songpengcheng Xia, Lizhou Lin et al.AAAI 2026 · 5 citations
- Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across ModalitiesChi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin et al.ACL 2026
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses Through Reasoning MLLMsZheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li et al.AAAI 2026 · 2 citations
- LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR UnderstandingSenqiao Yang, Jiaming Liu, Renrui Zhang, Mingjie Pan et al.AAAI 2025 · 17 citations
