Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration
Ui-Hyeop Shin, Jaehyun Ko, Woocheol Jeong, Hyung-Min Park
摘要
Speech restoration aims to recover clean speech from degraded recordings affected by noise, reverberation, bandwidth reduction, or other distortions, where input and output sampling rates may differ. Existing approaches typically assume matched input--output rates and apply redundant resampling, limiting native multi-rate processing. We formulate this gap as the extended sampling-frequency-independent (xSFI) setting, where a model must operate under decoupled input--output rates, and propose TF-Restormer, a query-based xSFI modeling framework. The model encodes only the observed input band and synthesizes the unobserved high-frequency band through extension queries with band-partitioned cross-attention, yielding an asymmetric encoder--decoder that allocates capacity to analysis while keeping synthesis lightweight. Trained with a perceptual loss, a scaled log-spectral loss, and adversarial supervision via an SFI-STFT discriminator, TF-Restormer attains balanced fidelity--perceptual quality as a single unified model, without redundant resampling across denoising, dereverberation, bandwidth extension, and combined-distortion benchmarks under multiple sampling rates.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsZeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan 等ICML 2024 · 被引用 341 次
- Siamese Masked AutoencodersAgrim Gupta, Jiajun Wu, Jia Deng, Fei-Fei LiNeurIPS 2023 · 被引用 113 次
相关 Paper
- Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image RestorationBin Ren, Runyi Yang, Qi Ma, Xu Zheng 等ICLR 2026
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- AdVerb: Visually Guided Audio DereverberationSanjoy Chowdhury, Sreyan Ghosh, Subhrajyoti Dasgupta, Anton Ratnarajah 等ICCV 2023 · 被引用 21 次
- RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value PairsZhouxia Wang, Jiawei Zhang, Runjian Chen, Wenping Wang 等CVPR 2022 · 被引用 109 次
- Lightweight Medical Image Restoration via Integrating Reliable Lesion-Semantic Driven PriorPengcheng Zheng, Kecheng Chen, Jiaxin Huang, Bohao Chen 等ACM MM 2025 · 被引用 1 次
