BAPEN: Towards Versatile Audio Phase Retrieval
Lingling Dai, Andong Li, Zhe Han, Chengshi Zheng, Xiaodong Li
摘要
Audio phase retrieval aims to reconstruct phase from the given magnitude and obtain the time-domain audio waveform. While deep learning techniques have promoted the development of this area, existing deep neural network (DNN)-based methods usually suffer from some inherent problems like limited generalization capability to different audio types, failing to adapt to different sampling rates, and inflexibility for varying computational complexity during the inference stage, which heavily hinder the development of the filed. To tackle these challenges, in this paper, we introduce a novel phase task estimation task called versatile auido phase retrieval and a Band-Aware Phase Estimation Network (BAPEN ) is proposed. Specifically, we first collect and establish a new benchmark for the task, which encompasses speech, sound effects, and music and the total duration is around 414 hours. Besides, a sub-band oriented framework is proposed, which involves hierarchical sub-band encoding/decoding and a dual-path network structure is specially devised for efficient narrow- and cross-band modeling, respectively. Furthermore, to enable dynamic control over the inference cost, we propose a simple yet effective sampling strategy for network depth augmentation during training. Both objective and subject results validate the promising performance of the BAPEN while possessing more flexible application ranges. Audio samples are available on: https://lingling-dai.github.io/BAPEN/.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- DegVoC: Revisiting Neural Vocoder from a Degradation PerspectiveAndong Li, Tong Lei, Lingling Dai, Kai Li 等AAAI 2026
- GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation TasksLingling Dai, Andong Li, Cheng Chi, Yifan Liang 等AAAI 2026
相关 Paper
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 被引用 387 次
- Complex-Cycle-Consistent Diffusion Model for Monaural Speech EnhancementYi Li, Yang Sun, Plamen P. AngelovAAAI 2025 · 被引用 2 次
- Enabling Fast and Universal Audio Adversarial Attack Using Generative ModelYi Xie, Zhuohang Li, Cong Shi, Jian Liu 等AAAI 2021 · 被引用 77 次
- Unsupervised Deep Learning for Phase Retrieval via Teacher-Student DistillationYuhui Quan, Zhile Chen, Tongyao Pang, Hui JiAAAI 2023 · 被引用 8 次
- DeAR: A Deep-Learning-Based Audio Re-recording Resilient WatermarkingChang Liu, Jie Zhang, Han Fang, Zehua Ma 等AAAI 2023 · 被引用 67 次
