Audio Super-Resolution with Latent Bridge Models
Chang Li, Zehua Chen, Liyuan Wang, Jun Zhu
摘要
Audio super-resolution (SR), i.e., upsampling the low-resolution (LR) waveform to the high-resolution (HR) version, has recently been explored with diffusion and bridge models, while previous methods often suffer from sub-optimal upsampling quality due to their uninformative generation prior. Towards high-quality audio super-resolution, we present a new system with latent bridge models (LBMs), where we compress the audio waveform into a continuous latent space and design an LBM to enable a latent-to-latent generation process that naturally matches the LR-to-HR upsampling process, thereby fully exploiting the instructive prior information contained in the LR waveform. To further enhance the training results despite the limited availability of HR samples, we introduce frequency-aware LBMs, where the prior and target frequency are taken as model input, enabling LBMs to explicitly learn an any-to-any upsampling process at the training stage. Furthermore, we design cascaded LBMs and present two prior augmentation strategies, where we make the first attempt to unlock the audio upsampling beyond 48 kHz and empower a seamless cascaded SR process, providing higher flexibility for audio post-production. Comprehensive experimental results evaluated on the VCTK, ESC-50, Song-Describer benchmark datasets and two internal testsets demonstrate that we achieve state-of-the-art objective and perceptual quality for any-to-48kHz SR across speech, audio, and music signals, as well as setting the first record for any-to-192kHz audio SR. Demo at https://AudioLBM.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Diffusion Bridge or Flow Matching? A Unifying Framework and Comparative AnalysisKaizhen Zhu, Mokai Pan, Zhechuan Yu, Jingya Wang 等ICML 2026 · 被引用 3 次
- GuidedBridge: Training-freely Improving Bridge Models with Prior GuidanceZehua Chen, Yucheng Yang, Binjie Yuan, Kaiwen Zheng 等ICML 2026
- PACE: Pretrained Audio Continual LearningChang Li, Kanglei Zhou, Liyuan WangICLR 2026
它引用的顶会 Paper37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- From Discrete Tokens to High-Fidelity Audio Using Multi-Band DiffusionRobin San Roman, Yossi Adi, Antoine Deleforge, Romain Serizel 等NeurIPS 2023 · 被引用 50 次
- FaithDiff: Unleashing Diffusion Priors for Faithful Image Super-resolutionJunyang Chen, Jinshan Pan, Jiangxin DongCVPR 2025
- MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video GenerationMingzhen Sun, Weining Wang, Yanyuan Qiao, Jiahui Sun 等ACM MM 2024 · 被引用 4 次
- SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech SynthesisYifan Liang, Andong Li, Kang Yang, Guochen Yu 等AAAI 2026
- Time-Frequency Domain Fusion Enhancement for Audio Super-ResolutionYe Tian, Zhe Wang, Jianguo Sun, Liguo ZhangACM MM 2024 · 被引用 1 次
