Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration
Ui-Hyeop Shin, Jaehyun Ko, Woocheol Jeong, Hyung-Min Park
Abstract
Speech restoration aims to recover clean speech from degraded recordings affected by noise, reverberation, bandwidth reduction, or other distortions, where input and output sampling rates may differ. Existing approaches typically assume matched input--output rates and apply redundant resampling, limiting native multi-rate processing. We formulate this gap as the extended sampling-frequency-independent (xSFI) setting, where a model must operate under decoupled input--output rates, and propose TF-Restormer, a query-based xSFI modeling framework. The model encodes only the observed input band and synthesizes the unobserved high-frequency band through extension queries with band-partitioned cross-attention, yielding an asymmetric encoder--decoder that allocates capacity to analysis while keeping synthesis lightweight. Trained with a perceptual loss, a scaled log-spectral loss, and adversarial supervision via an SFI-STFT discriminator, TF-Restormer attains balanced fidelity--perceptual quality as a single unified model, without redundant resampling across denoising, dereverberation, bandwidth extension, and combined-distortion benchmarks under multiple sampling rates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsZeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan et al.ICML 2024 · 341 citations
- Siamese Masked AutoencodersAgrim Gupta, Jiajun Wu, Jia Deng, Fei-Fei LiNeurIPS 2023 · 113 citations
Related papers
- Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image RestorationBin Ren, Runyi Yang, Qi Ma, Xu Zheng et al.ICLR 2026
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- AdVerb: Visually Guided Audio DereverberationSanjoy Chowdhury, Sreyan Ghosh, Subhrajyoti Dasgupta, Anton Ratnarajah et al.ICCV 2023 · 21 citations
- RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value PairsZhouxia Wang, Jiawei Zhang, Runjian Chen, Wenping Wang et al.CVPR 2022 · 109 citations
- Lightweight Medical Image Restoration via Integrating Reliable Lesion-Semantic Driven PriorPengcheng Zheng, Kecheng Chen, Jiaxin Huang, Bohao Chen et al.ACM MM 2025 · 1 citation
