MoXaRt: Audio-Visual Object-Guided Sound Interaction for XR
Tianyu Xu, Sieun Kim, Qianhui Zheng, Ruoyu Xu, Tejasvi Ravi, Anuva Kulkarni, Katrina Passarella-Ward, Junyi Zhu, Adarsh Kowdle
摘要
In Extended Reality (XR), complex acoustic environments often overwhelm users, compromising both scene awareness and social engagement due to entangled sound sources. We introduce MoXaRt, a real-time XR system that uses audio-visual cues to separate these sources and enable fine-grained sound interaction. MoXaRt’s core is a cascaded architecture that performs coarse, audio-only separation in parallel with visual detection of sources (e.g., faces, instruments). These visual anchors then guide refinement networks to isolate individual sources, separating complex mixes of up to 5 concurrent sources (e.g., 2 voices + 3 instruments) with ∼ 2 second processing latency. We validate MoXaRt through a technical evaluation on a new dataset of 30 one-minute recordings featuring concurrent speech and music, and a 22-participant user study. Empirical results indicate that our system significantly enhances speech intelligibility, yielding a 36.2% (p < 0.01) increase in listening comprehension within adversarial acoustic environments while substantially reducing cognitive load (p < 0.001), thereby paving the way for more perceptive and socially adept XR experiences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss 等NeurIPS 2020 · 被引用 227 次
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey 等ICLR 2021 · 被引用 83 次
- GesturAR: An Authoring System for Creating Freehand Interactive Augmented Reality ApplicationsTianyi Wang, Xun Qian, Fengming He, Xiyun Hu 等UIST 2021 · 被引用 75 次
- ProtoSound: A Personalized and Scalable Sound Recognition System for Deaf and Hard-of-Hearing UsersDhruv Jain, Khoa Huynh Anh Nguyen, Steven M. Goodman, Rachel Grossman-Kahn 等CHI 2022 · 被引用 45 次
相关 Paper
- Auptimize: Optimal Placement of Spatial Audio Cues for Extended RealityHyunsung Cho, Alexander Wang, Divya Kartik, Emily Liying Xie 等UIST 2024 · 被引用 19 次
- Perceptually-Guided Acoustic "Foveation"Xi Peng, Kenneth Chen, Iran Roman, Juan Pablo Bello 等IEEE VR 2025 · 被引用 2 次
- Environment Spatial Restitution for Remote Physical AR CollaborationBruno Caby, Guillaume Bataille, Florence Danglade, Jean-Rémy ChardonnetIEEE VR 2025 · 被引用 1 次
- Overcoming Translation Delays: Towards Better Subtitle Design for Foreign Language Conversations in Extended RealityZiming Li, Rongkai Shi, Hongji Li, Jialin Wang 等CHI 2026 · 被引用 1 次
- ECHO: Efficient Head-Orientation-Guided Real-Time Sound Spatialization for Virtual RealityHaiyu Wang, Tianhua Xia, Sai Qian ZhangISCA 2026
