AV-Cloud: Spatial Audio Rendering Through Audio-Visual Cloud Splatting
Mingfei Chen, Eli Shlizerman
摘要
We propose a novel approach for rendering high-quality spatial audio for 3D scenes that is in synchrony with the visual stream but does not rely or explicitly conditioned on the visual rendering. We demonstrate that such an approach enables the experience of immersive virtual tourism - performing a real-time dynamic navigation within the scene, experiencing both audio and visual content. Current audio-visual rendering approaches typically rely on visual cues, such as images, and thus visual artifacts could cause inconsistency in the audio quality. Furthermore, when such approaches are incorporated with visual rendering, audio generation at each viewpoint occurs after the rendering of the image of the viewpoint and thus could lead to audio lag that affects the integration of audio and visual streams. Our proposed approach, AV-Cloud , overcomes these challenges by learning the representation of the audio-visual scene based on a set of sparse AV anchor points, that constitute the Audio-Visual Cloud, and are derived from the camera calibration. The Audio-Visual Cloud serves as an audio-visual representation from which the generation of spatial audio for arbitrary listener location can be generated. In particular, we propose a novel module Audio-Visual Cloud Splatting which decodes AV anchor points into a spatial audio transfer function for the arbitrary viewpoint of the target listener. This function, applied through the Spatial Audio Render Head module, transforms monaural input into viewpoint-specific spatial audio. As a result, AV-Cloud efficiently renders the spatial audio aligned with any visual viewpoint and eliminates the need for pre-rendered images. We show that AV-Cloud surpasses current state-of-the-art accuracy on audio reconstruction, perceptive quality, and acoustic effects on two real-world datasets. AV-Cloud also outperforms previous methods when tested on scenes “in the wild”.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- -AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?Susan Liang, Chao Huang, Yunlong Tang, Zeliang Zhang 等ICCV 2025 · 被引用 4 次
- Resounding Acoustic Fields with ReciprocityZitong Lan, Yiduo Hao, Mingmin ZhaoNeurIPS 2025 · 被引用 3 次
- Few-shot Acoustic Synthesis with Multimodal Flow MatchingAmandine BrunettoCVPR 2026 · 被引用 2 次
- Hearing Anywhere in Any EnvironmentXiulong Liu, Anurag Kumar, Paul Calamia, Sebastià Vicenc Amengual Garí 等CVPR 2025
它引用的顶会 Paper15
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- SparseNeRF: Distilling Depth Ranking for Few-shot Novel View SynthesisGuangcong Wang, Zhaoxi Chen, Chen Change Loy, Ziwei LiuICCV 2023 · 被引用 309 次
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum 等NeurIPS 2022 · 被引用 153 次
- ADOP: approximate differentiable one-pixel point renderingDarius Rückert, Linus Franke, Marc StammingerSIGGRAPH 2022 · 被引用 127 次
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 被引用 92 次
相关 Paper
- AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic SynthesisSwapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng 等NeurIPS 2024 · 被引用 22 次
- Be Everywhere - Hear Everything (BEE): Audio Scene Reconstruction by Sparse Audio-Visual SamplesMingfei Chen, Kun Su, Eli ShlizermanICCV 2023 · 被引用 13 次
- Sonic4D: Spatial Audio Generation for Immersive 4D Scene ExplorationSiyi Xie, Hanxin Zhu, Xinyi Chen, Tianyu He 等AAAI 2026 · 被引用 4 次
- AudioEar: Single-View Ear Reconstruction for Personalized Spatial AudioXiaoyang Huang, Yanjun Wang, Yang Liu, Bingbing Ni 等AAAI 2023 · 被引用 5 次
- SonoWorld: From One Image to a 3D Audio-Visual SceneDerong Jin, Xiyi Chen, Ming C. Lin, Ruohan GaoCVPR 2026 · 被引用 4 次
