MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
Tianyu Xu, Sieun Kim, Qianhui Zheng, Ruoyu Xu, Tejasvi Ravi, Anuva Kulkarni, Katrina Passarella-Ward, Junyi Zhu, Adarsh Kowdle
Abstract
In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce MoXaRt, a real-time XR system that uses audio-visual cues to separate these sources and enable fine-grained sound interaction. MoXaRt’s core is a cascaded architecture that performs coarse, audio-only separation in parallel with visual detection of sources (e.g., faces, instruments). These visual anchors then guide refinement networks to isolate individual sources, separating complex mixes of up to 5 concurrent sources (e.g., 2 voices + 3 instruments) with ∼ 2 second processing latency. We validate MoXaRt through a technical evaluation on a new dataset of 30 one-minute recordings featuring concurrent speech and music, and a 22-participant user study. Empirical results indicate that our system significantly enhances speech intelligibility, yielding a 36.2% (p < 0.01) increase in listening comprehension within adversarial acoustic environments while substantially reducing cognitive load (p < 0.001), thereby paving the way for more perceptive and socially adept XR experiences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f8b8464-fa2c-4e25-a299-12c465a2c58cBuilds on18
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey et al.ICLR 2021 · 83 citations
- GesturAR: An Authoring System for Creating Freehand Interactive Augmented Reality ApplicationsTianyi Wang, Xun Qian, Fengming He, Xiyun Hu et al.UIST 2021 · 75 citations
- ProtoSound: A Personalized and Scalable Sound Recognition System for Deaf and Hard-of-Hearing UsersDhruv Jain, Khoa Huynh Anh Nguyen, Steven M. Goodman, Rachel Grossman-Kahn et al.CHI 2022 · 45 citations
Related papers
- Auptimize: Optimal Placement of Spatial Audio Cues for Extended RealityHyunsung Cho, Alexander Wang, Divya Kartik, Emily Liying Xie et al.UIST 2024 · 19 citations
- Perceptually-Guided Acoustic "Foveation"Xi Peng, Kenneth Chen, Iran Roman, Juan Pablo Bello et al.IEEE VR 2025 · 2 citations
- Environment Spatial Restitution for Remote Physical AR CollaborationBruno Caby, Guillaume Bataille, Florence Danglade, Jean-Rémy ChardonnetIEEE VR 2025 · 1 citation
- Overcoming Translation Delays: Towards Better Subtitle Design for Foreign Language Conversations in Extended RealityZiming Li, Rongkai Shi, Hongji Li, Jialin Wang et al.CHI 2026 · 1 citation
- ECHO: Efficient Head-Orientation-Guided Real-Time Sound Spatialization for Virtual RealityHaiyu Wang, Tianhua Xia, Sai Qian ZhangISCA 2026
