SpeechOp: Inference-Time Task Composition for Generative Speech Processing
Justin Lovelace, Rithesh Kumar, Jiaqi Su, Ke Chen, Kilian Q. Weinberger, Zeyu Jin
摘要
While generative Text-to-Speech (TTS) systems leverage vast "in-the-wild" data to achieve remarkable success, speech-to-speech processing tasks like enhancement face data limitations, which lead data-hungry generative approaches to distort speech content and speaker identity. To bridge this gap, we present SpeechOp, a multi-task latent diffusion model that transforms pre-trained TTS models into a universal speech processor capable of performing a wide range of speech tasks and composing them in novel ways at inference time. By adapting a pre-trained TTS model, SpeechOp inherits a rich understanding of natural speech, accelerating training and improving S2S task quality, while simultaneously enhancing core TTS performance. Finally, we introduce Implicit Task Composition (ITC), a novel pipeline where ASR-derived transcripts (e.g., from Whisper) guide SpeechOp's enhancement via our principled inference-time task composition. ITC achieves state-of-the-art content preservation by robustly combining web-scale speech understanding with SpeechOp's generative capabilities. Audio samples are available at https://justinlovelace.github.io/projects/speechop .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
相关 Paper
- Generative Pre-training for Speech with Flow MatchingAlexander H. Liu, Matthew Le, Apoorv Vyas, Bowen Shi 等ICLR 2024 · 被引用 66 次
- From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech RecognitionTianduo Wang, Lu Xu, Wei Lu, Shanbo ChengEMNLP 2025 · 被引用 1 次
- DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific FactorsKeon Lee, Dong Won Kim, Jaehyeon Kim, Seungjun Chung 等ICLR 2025
- Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier GuidanceHeeseung Kim, Sungwon Kim, Sungroh YoonICML 2022 · 被引用 133 次
- Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?Qingkai Fang, Shaolei Zhang, Zhengrui Ma, Min Zhang 等ACL 2024 · 被引用 1 次
