Listening to Sounds of Silence for Speech Denoising
Ruilin Xu, Rundi Wu, Yuko Ishiwaka, Carl Vondrick, Changxi Zheng
摘要
We introduce a deep learning model for speech denoising, a long-standing challenge in audio analysis arising in numerous applications. Our approach is based on a key observation about human speech: there is often a short pause between each sentence or word. In a recorded speech signal, those pauses introduce a series of time periods during which only noise is present. We leverage these incidental silent intervals to learn a model for automatic speech denoising given only mono-channel audio. Detected silent intervals over time expose not just pure noise but its time-varying features, allowing the model to learn noise dynamics and suppress it from the speech signal. Experiments on multiple datasets confirm the pivotal role of silent interval detection for speech denoising, and our method outperforms several state-of-the-art denoising methods, including those that accept only audio input (like ours) and those that denoise based on audiovisual input (and hence require more information). We also show that our method enjoys excellent generalization properties, such as denoising spoken languages not seen during training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Real-Time Neural Voice CamouflageMia Chiquier, Chengzhi Mao, Carl VondrickICLR 2022 · 被引用 9 次
- Un-Rocking Drones: Foundations of Acoustic Injection Attacks and Recovery ThereofJinseob Jeong, Dongkwan Kim, Joon-Ha Jang, Juhwan Noh 等NDSS 2023
- I Can Hear You: Selective Robust Training for Deepfake Audio DetectionZirui Zhang, Wei Hao, Aroon Sankoh, William Lin 等ICLR 2025
- D4AM: A General Denoising Framework for Downstream Acoustic ModelsChi-Chang Lee, Yu Tsao, Hsin-Min Wang, Chu-Song ChenICLR 2023
它引用的顶会 Paper2
相关 Paper
- Unsupervised Deep Video Denoising with Untrained NetworkHuan Zheng, Tongyao Pang, Hui JiAAAI 2023 · 被引用 14 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- Circumventing Shortcuts in Audio-visual Deepfake Detection Datasets with Unsupervised LearningStefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta OneataCVPR 2025
- Towards Accurate Lip-to-Speech Synthesis in-the-WildSindhu B. Hegde, Rudrabha Mukhopadhyay, C. V. Jawahar, Vinay P. NamboodiriACM MM 2023 · 被引用 9 次
- Watch Your Mouth: Silent Speech Recognition with Depth SensingXue Wang, Zixiong Su, Jun Rekimoto, Yang ZhangCHI 2024 · 被引用 19 次
