PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models
Jiangong Chen, Mingyu Zhu, Bin Li
Abstract
Multimodal Large Language Models (MLLMs) enhance collaboration in Extended Reality (XR) environments by enabling flexible object and animation creation through the combination of natural language and visual inputs. However, visual data captured by XR headsets includes real-world backgrounds that may contain irrelevant or sensitive user information, such as credit cards left on the table or facial identities of other users. Uploading those frames to cloud-based MLLMs poses serious privacy risks, particularly when such data is processed without explicit user consent. Additionally, existing colocation and synchronization mechanisms in commercial XR APIs rely on time-consuming, privacy-invasive environment scanning and struggle to adapt to the highly dynamic nature of MLLM-integrated XR environments. In this paper, we propose PRISM-XR, a novel framework that facilitates multi-user collaboration in XR by providing privacy-aware MLLM integration. PRISM-XR employs intelligent frame preprocessing on the edge server to filter sensitive data and remove irrelevant context before communicating with cloud generative AI models. Additionally, we introduce a lightweight registration process and a fully customizable content-sharing mechanism to enable efficient, accurate, and privacy-preserving content synchronization among users. Our numerical evaluation results indicate that the proposed platform achieves nearly 90% accuracy in fulfilling user requests and less than 0.27 seconds registration time while maintaining spatial inconsistencies of less than 3.5 cm. Furthermore, we conducted an IRB-approved user study with 28 participants, demonstrating that our system could automatically filter highly sensitive objects in over 90% of scenarios while maintaining strong overall usability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8614d860-470c-482b-b85a-243e6b49daaaBuilds on16
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey et al.CHI 2024 · 124 citations
- Privacy-Enhancing Technology and Everyday Augmented Reality: Understanding Bystanders' Varying Needs for Awareness and ConsentJoseph O'Hagan, Pejman Saeghe, Jan Gugenheimer, Daniel Medeiros et al.UbiComp 2023 · 102 citations
Related papers
- Explainable XR: Understanding User Behaviors of XR Environments Using LLM-Assisted Analytics FrameworkYoonsang Kim, Zainab Aamir, Mithilesh Kumar Singh, Saeed Boorboor et al.IEEE VR 2025 · 27 citations
- PRISM: Privacy-Aware Routing for Adaptive Cloud-Edge LLM Inference via Semantic Sketch CollaborationJunfei Zhan, Haoxun Shen, Zheng Lin, Tengjiao HeAAAI 2026 · 4 citations
- LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language ModelsJiangong Chen, Xiaoyi Wu, Tian Lan, Bin LiIEEE VR 2025 · 20 citations
- EmBARDiment: an Embodied AI Agent for Productivity in XRRiccardo Bovo, Steven Abreu, Karan Ahuja, Eric J. Gonzalez et al.IEEE VR 2025 · 19 citations
- Boundary Probing for Input Privacy Protection when Using LMM ServicesXiaofei Hui, Haoxuan Qu, Ping Hu, Hossein Rahmani et al.ICCV 2025
