PADS-TAL: Padding-Annealed Diffusion Sampling in Text-Aware Latent Space for Robust and Diverse Text-to-Music Generation
Taekoan Yoo, Wonkyung Jung, Kyunghun Kim, Kyeongbo Kong
Abstract
Text-to-Music diffusion models are increasingly used in real-world applications, yet deployment remains challenging: generations can collapse to limited patterns even with diverse initial noise and prompts, and inference-time diversity control often harms text alignment and fidelity by distorting key prompt cues established in early denoising. To address this, we propose Padding-Annealed Diffusion Sampling, which perturbs only a padding-indexed subspace while keeping non-padding conditioning fixed, enabling controlled exploration with reduced semantic drift. However, in a text-unaware VAE latent space, such exploration is less likely to stay within genre-faithful neighborhoods, limiting genre-consistent diversity. We therefore introduce Text-Aware Latent space that aligns local neighborhoods with text-implied genre structure, promoting genre-consistent diversity. Together, the two techniques form a unified pipeline that, compared to prior methods that perturb the full conditioning, achieves a better text alignment-diversity trade-off: at comparable text alignment, it delivers 15.4% higher diversity with a relatively small fidelity drop, and further improves within-genre diversity by 71.6%. Generated samples are available at https://pads-tal.github.io/PADS-TAL. io * Equal contribution -Taekoan Yoo led the project, with substantial contributions from Kyunghun Kim and Wonkyung Jung.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
Related papers
- Diverse Text-to-Image Generation via Contrastive Noise OptimizationByungjun Kim, Soobin Um, Jong Chul YeICLR 2026 · 11 citations
- Efficient Neural Music GenerationMax W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin et al.NeurIPS 2023 · 95 citations
- DITTO: Diffusion Inference-Time T-Optimization for Music GenerationZachary Novack, Julian J. McAuley, Taylor Berg-Kirkpatrick, Nicholas J. BryanICML 2024 · 81 citations
- CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed SamplingSeyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges et al.ICLR 2024 · 115 citations
- Be Decisive: Noise-Induced Layouts for Multi-Subject GenerationOmer Dahary, Yehonathan Cohen, Or Patashnik, Kfir Aberman et al.SIGGRAPH 2025 · 3 citations
