MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modeling
Yusong Wu, Ethan Manilow, Yi Deng, Rigel Swavely, Kyle Kastner, Tim Cooijmans, Aaron C. Courville, Cheng-Zhi Anna Huang, Jesse H. Engel
Abstract
Musical expression requires control of both what notes are played, and how they are performed. Conventional audio synthesizers provide detailed expressive controls, but at the cost of realism. Black-box neural audio synthesis and concatenative samplers can produce realistic audio, but have few mechanisms for control. In this work, we introduce MIDI-DDSP a hierarchical model of musical instruments that enables both realistic neural audio synthesis and detailed user control. Starting from interpretable Differentiable Digital Signal Processing (DDSP) synthesis parameters, we infer musical notes and high-level properties of their expressive performance (such as timbre, vibrato, dynamics, and articulation). This creates a 3-level hierarchy (notes, performance, synthesis) that affords individuals the option to intervene at each level, or utilize trained priors (performance given notes, synthesis given performance) for creative assistance. Through quantitative experiments and listening tests, we demonstrate that this hierarchy can reconstruct high-fidelity audio, accurately predict performance attributes for a note sequence, independently manipulate the attributes of a given performance, and as a complete system, generate realistic audio from a novel note sequence. By utilizing an interpretable hierarchy, with multiple levels of granularity, MIDI-DDSP opens the door to assistive tools to empower individuals across a diverse range of musical experience. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a10fb4f-c567-4ddf-88a7-7a3e3b247b7dCited by top-tier papers5
- MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music CompositionPhilippe Pasquier, Jeff Ens, Nathan Fradet, Paul Triana et al.AAAI 2025 · 14 citations
- Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion SimulationJin Woo Lee, Jaehyun Park, Min Jun Choi, Kyogu LeeNeurIPS 2024 · 9 citations
- MID-FiLD: MIDI Dataset for Fine-Level DynamicsJesung Ryu, Seungyeon Rhyu, Hong-Gyu Yoon, Eunchong Kim et al.AAAI 2024 · 4 citations
- Detecting Music Performance Errors with TransformersBenjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos, Tim Nadolsky et al.AAAI 2025 · 3 citations
- LadderSym: A Multimodal Interleaved Transformer for Music Practice Error DetectionBenjamin Shiue-Hal Chou, Purvish Jajal, Nicholas Eliopoulos, James C. Davis et al.ICLR 2026 · 2 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
Related papers
- DDSP: Differentiable Digital Signal ProcessingJesse H. Engel, Lamtharn Hantrakul, Chenjie Gu, Adam RobertsICLR 2020 · 467 citations
- Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic MusicHongju Su, Ke Li, Lan Yang, Honggang Zhang et al.ACL 2026
- MUSIC: Learning Muscle-Driven Dexterous Hand ControlPei Xu, Yufei Ye, Shuchun Sun, Yu Ding et al.SIGGRAPH 2026
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-TrainingHong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia et al.ICML 2026 · 3 citations
- DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference TasksXutong Jin, Chenxi Xu, Ruohan Gao, Jiajun Wu et al.SIGGRAPH 2024 · 3 citations
