BAPEN: Towards Versatile Audio Phase Retrieval
Lingling Dai, Andong Li, Zhe Han, Chengshi Zheng, Xiaodong Li
Abstract
Audio phase retrieval aims to reconstruct phase from the given magnitude and obtain the time-domain audio waveform. While deep learning techniques have promoted the development of this area, existing deep neural network (DNN)-based methods usually suffer from some inherent problems like limited generalization capability to different audio types, failing to adapt to different sampling rates, and inflexibility for varying computational complexity during the inference stage, which heavily hinder the development of the filed. To tackle these challenges, in this paper, we introduce a novel phase task estimation task called versatile auido phase retrieval and a Band-Aware Phase Estimation Network (BAPEN ) is proposed. Specifically, we first collect and establish a new benchmark for the task, which encompasses speech, sound effects, and music and the total duration is around 414 hours. Besides, a sub-band oriented framework is proposed, which involves hierarchical sub-band encoding/decoding and a dual-path network structure is specially devised for efficient narrow- and cross-band modeling, respectively. Furthermore, to enable dynamic control over the inference cost, we propose a simple yet effective sampling strategy for network depth augmentation during training. Both objective and subject results validate the promising performance of the BAPEN while possessing more flexible application ranges. Audio samples are available on: https://lingling-dai.github.io/BAPEN/.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- DegVoC: Revisiting Neural Vocoder from a Degradation PerspectiveAndong Li, Tong Lei, Lingling Dai, Kai Li et al.AAAI 2026
- GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation TasksLingling Dai, Andong Li, Cheng Chi, Yifan Liang et al.AAAI 2026
Related papers
- PHASEN: A Phase-and-Harmonics-Aware Speech Enhancement NetworkDacheng Yin, Chong Luo, Zhiwei Xiong, Wenjun ZengAAAI 2020 · 387 citations
- Complex-Cycle-Consistent Diffusion Model for Monaural Speech EnhancementYi Li, Yang Sun, Plamen P. AngelovAAAI 2025 · 2 citations
- Enabling Fast and Universal Audio Adversarial Attack Using Generative ModelYi Xie, Zhuohang Li, Cong Shi, Jian Liu et al.AAAI 2021 · 77 citations
- Unsupervised Deep Learning for Phase Retrieval via Teacher-Student DistillationYuhui Quan, Zhile Chen, Tongyao Pang, Hui JiAAAI 2023 · 8 citations
- DeAR: A Deep-Learning-Based Audio Re-recording Resilient WatermarkingChang Liu, Jie Zhang, Han Fang, Zehua Ma et al.AAAI 2023 · 67 citations
