AudioEar: Single-View Ear Reconstruction for Personalized Spatial Audio
Xiaoyang Huang, Yanjun Wang, Yang Liu, Bingbing Ni, Wenjun Zhang, Jinxian Liu, Teng Li
摘要
Spatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source positions. In this work, we address this problem from an interdisciplinary perspective. The rendering of spatial audio is strongly correlated with the 3D shape of human bodies, particularly ears. To this end, we propose to achieve personalized spatial audio by reconstructing 3D human ears with single-view images. First, to benchmark the ear reconstruction task, we introduce AudioEar3D, a high-quality 3D ear dataset consisting of 112 point cloud ear scans with RGB images. To self-supervisedly train a reconstruction model, we further collect a 2D ear dataset composed of 2,000 images, each one with manual annotation of occlusion and 55 landmarks, named AudioEar2D. To our knowledge, both datasets have the largest scale and best quality of their kinds for public use. Further, we propose AudioEarM, a reconstruction method guided by a depth estimation network that is trained on synthetic data, with two loss functions tailored for ear data. Lastly, to fill the gap between the vision and acoustics community, we develop a pipeline to integrate the reconstructed ear mesh with an off-the-shelf 3D human body and simulate a personalized Head-Related Transfer Function (HRTF), which is the core of spatial audio rendering. Code and data are publicly available in https://github.com/seanywang0408/AudioEar.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 被引用 662 次
- Image GANs meet Differentiable Rendering for Inverse Graphics and Interpretable 3D Neural RenderingYuxuan Zhang, Wenzheng Chen, Huan Ling, Jun Gao 等ICLR 2021 · 被引用 140 次
- Boosting Point Clouds Rendering via Radiance MappingXiaoyang Huang, Yi Zhang, Bingbing Ni, Teng Li 等AAAI 2023 · 被引用 15 次
- Representation-Agnostic Shape FieldsXiaoyang Huang, Jiancheng Yang, Yanjun Wang, Ziyu Chen 等ICLR 2022 · 被引用 7 次
- High-Fidelity Neural Human Motion Transfer From Monocular VideoMoritz Kappel, Vladislav Golyanik, Mohamed A. Elgharib, Jann-Ole Henningson 等CVPR 2021
相关 Paper
- Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and AudioXudong Xu, Dejan Markovic, Jacob Sandakly, Todd Keebler 等NeurIPS 2023 · 被引用 9 次
- HRTF Estimation in the WildVivek Jayaram, Ira Kemelmacher-Shlizerman, Steven M. SeitzUIST 2023 · 被引用 9 次
- Graph Neural Field with Spatial-Correlation Augmentation for HRTF PersonalizationDe Hu, Junsheng Hu, Cuicui JiangAAAI 2026
- AV-Cloud: Spatial Audio Rendering Through Audio-Visual Cloud SplattingMingfei Chen, Eli ShlizermanNeurIPS 2024 · 被引用 14 次
- Personalizing head related transfer functions for earablesZhijian Yang, Romit Roy ChoudhurySIGCOMM 2021 · 被引用 26 次
