SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
Koichi Saito, Dongjun Kim, Takashi Shibuya, Chieh-Hsin Lai, Zhi Zhong, Yuhta Takida, Yuki Mitsufuji
摘要
Recent high-quality diffusion-based sound generation models can serve as valuable tools for sound content creators. However, despite producing high-quality sounds, these models often suffer from slow inference speeds. This drawback burdens creators, who typically refine their sounds through trial and error to align sounds with their artistic intentions. To address this issue, we introduce Sound Consistency Trajectory Models (SoundCTM). Our model enables flexible transitioning between high-quality 1-step sound generation and superior sound quality through multi-step generation. This allows creators to initially control sounds with 1-step samples before refining them through multi-step generation. We reframe original CTM's training framework and introduce a novel feature distance by utilizing the teacher's network for a distillation loss. Additionally, while distilling classifier-free guided trajectories, we train conditional and unconditional student models simultaneously and interpolate between these models during inference. SoundCTM achieves both promising 1-step and multi-step real-time sound generation. Audio samples are available at https://anonymus-soundctm.github.io/soundctm/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean FlowsXiquan Li, Junxi Liu, Yuzhe Liang, Zhikang Niu 等ACL 2026 · 被引用 25 次
- SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music EditingXinlei Niu, Kin Wai Cheuk, Jing Zhang, Naoki Murata 等AAAI 2026 · 被引用 5 次
- Sounding that Object: Interactive Object-Aware Image to Audio GenerationTingle Li, Baihe Huang, Xiaobin Zhuang, Dongya Jia 等ICML 2025
- BNMusic: Blending Environmental Noises into Personalized MusicChi Zuo, Martin Bo Møller, Pablo Martínez-Nuevo, Huayang Huang 等NeurIPS 2025
它引用的顶会 Paper17
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
相关 Paper
- Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of DiffusionDongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Naoki Murata 等ICLR 2024 · 被引用 377 次
- Theory of Consistency Diffusion Models: Distribution Estimation Meets Fast SamplingZehao Dou, Minshuo Chen, Mengdi Wang, Zhuoran YangICML 2024 · 被引用 11 次
- AudioLCM: Efficient and High-Quality Text-to-Audio Generation with Minimal Inference StepsHuadai Liu, Rongjie Huang, Yang Liu, Hengyuan Cao 等ACM MM 2024 · 被引用 4 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- Consistency Trajectory Matching for One-Step Generative Super-ResolutionWeiyi You, Mingyang Zhang, Leheng Zhang, Xingyu Zhou 等ICCV 2025 · 被引用 5 次
