DDSP: Differentiable Digital Signal Processing
Jesse H. Engel, Lamtharn Hantrakul, Chenjie Gu, Adam Roberts
Abstract
Most generative models of audio directly generate samples in one of two domains: time or frequency. While sufficient to express any signal, these representations are inefficient, as they do not utilize existing knowledge of how sound is generated and perceived. A third approach (vocoders/synthesizers) successfully incorporates strong domain knowledge of signal processing and perception, but has been less actively researched due to limited expressivity and difficulty integrating with modern auto-differentiation-based machine learning methods. In this paper, we introduce the Differentiable Digital Signal Processing (DDSP) library, which enables direct integration of classic signal processing elements with deep learning methods. Focusing on audio synthesis, we achieve high-fidelity generation without the need for large autoregressive models or adversarial losses, demonstrating that DDSP enables utilizing strong inductive biases without losing the expressive power of neural networks. Further, we show that combining interpretable modules permits manipulation of each separate model component, with applications such as independent control of pitch and loudness, realistic extrapolation to pitches not seen during training, blind dereverberation of room acoustics, transfer of extracted room acoustics to new environments, and transformation of timbre between disparate sources. In short, DDSP enables an interpretable and modular approach to generative modeling, without sacrificing the benefits of deep learning. The library will be made available upon paper acceptance and we encourage further contributions from the community and domain experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised RepresentationsHyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee et al.NeurIPS 2021 · 200 citations
- MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical ModelingYusong Wu, Ethan Manilow, Yi Deng, Rigel Swavely et al.ICLR 2022 · 65 citations
- WaveGrad: Estimating Gradients for Waveform GenerationNanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss et al.ICLR 2021 · 44 citations
- MusicRL: Aligning Music Generation to Human PreferencesGeoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent et al.ICML 2024 · 41 citations
- End-to-end Adversarial Text-to-SpeechJeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen et al.ICLR 2021 · 33 citations
Related papers
- SCRAPL: Scattering Transform with Random Paths for Machine LearningChristopher Mitcheltree, Vincent Lostanlen, Emmanouil Benetos, Mathieu LagrangeICLR 2026
- From Discrete Tokens to High-Fidelity Audio Using Multi-Band DiffusionRobin San Roman, Yossi Adi, Antoine Deleforge, Romain Serizel et al.NeurIPS 2023 · 50 citations
- Hearing Anything AnywhereMason Long Wang, Ryosuke Sawata, Samuel Clarke, Ruohan Gao et al.CVPR 2024 · 6 citations
- Differentiable Geometric Acoustic Path Tracing using Time-Resolved Path Replay BackpropagationUgo Paavo Finnendahl, Markus Worchel, Tobias Jüterbock, Daniel Wujecki et al.SIGGRAPH 2025 · 3 citations
- Deep Neural Room Acoustics PrimitiveYuhang He, Anoop Cherian, Gordon Wichern, Andrew MarkhamICML 2024 · 6 citations
