A Multimodal Evaluation Framework for Spatial Audio Playback Systems: From Localization to Listener Preference
Changhao Pan, Wenxiang Guo, Yu Zhang, Zhiyuan Zhu, Zhetao Chen, Han Wang, Zhou Zhao
摘要
Spatial audio playback defines immersive listening. However, objective evaluation methods for perceptual dimensions like sound field and sound image remain underdeveloped, hindered by the lack of fine-grained spatial audio datasets and the neglect of echoes and reverberation in diverse playback conditions. To address these challenges, we propose MESA, a multi-modal evaluation framework for spatial audio systems, and introduce PSA-MOS, a high-quality multi-scene spatial audio dataset. Specifically: 1) PSA-MOS provides 50 hours of high-quality spatial audio recordings spanning 6 playback scenarios and 7 device types, with detailed localization annotations and fine-grained MOS ratings across four perceptual dimensions. 2) We develop SAE-Encoder, a spatial audio encoder that captures both acoustic-spatial cues and fine-grained perceptual patterns. 3) MESA integrates visual scene context to enhance evaluation robustness through echo and reverberation modeling. Experimental results demonstrate that SAE-Encoder achieves superior performance in SELD tasks. With a two-stage training strategy, MESA exhibits strong correlation with human perceptual assessments, effectively guiding spatial audio quality optimization. The demos are available at https://david-pigeon.github.io/mesaDemo.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- OWL : Geometry-Aware Spatial Reasoning for Audio Large Language ModelsSubrata Biswas, Mohammad Nur Hossain Khan, Bashima IslamICLR 2026 · 被引用 9 次
- 🎧MOSPA: Human Motion Generation Driven by Spatial AudioShuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan 等NeurIPS 2025 · 被引用 13 次
- PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape MappingSubash Khanal, Eric Xing, Srikumar Sastry, Aayush Dhakal 等ACM MM 2024 · 被引用 3 次
- AudioEar: Single-View Ear Reconstruction for Personalized Spatial AudioXiaoyang Huang, Yanjun Wang, Yang Liu, Bingbing Ni 等AAAI 2023 · 被引用 5 次
- In-the-wild Audio Spatialization with Flexible Text-guided LocalizationTianrui Pan, Jie Liu, Zewen Huang, Jie Tang 等ACL 2025 · 被引用 2 次
