Test-time Alignment of Diffusion Models without Reward Over-optimization
Sunwoo Kim, Minkyu Kim, Dongmin Park
摘要
Diffusion models excel in generative tasks, but aligning them with specific objectives while maintaining their versatility remains challenging. Existing fine-tuning methods often suffer from reward over-optimization, while approximate guidance approaches fail to optimize target rewards effectively. Addressing these limitations, we propose a training-free, test-time method based on Sequential Monte Carlo (SMC) to sample from the reward-aligned target distribution. Our approach, tailored for diffusion sampling and incorporating tempering techniques, achieves comparable or superior target rewards to fine-tuning methods while preserving diversity and cross-reward generalization. We demonstrate its effectiveness in single-reward optimization, multi-objective scenarios, and online black-box optimization. This work offers a robust solution for aligning diffusion models with diverse downstream objectives without compromising their general capabilities. Code is available at https://github.com/krafton-ai/DAS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Derivative-Free Guidance in Continuous and Discrete Diffusion Models with Soft Value-based DecodingXiner Li, Yulai Zhao, Chenyu Wang, Gabriele Scalia 等NeurIPS 2025 · 被引用 147 次
- Reasoning with Sampling: Your Base Model is Smarter Than You ThinkAayush Karan, Yilun DuICLR 2026 · 被引用 87 次
- Inference-time scaling of diffusion models through classical searchXiangcheng Zhang, Haowei Lin, Haotian Ye, James Y. Zou 等ICLR 2026 · 被引用 57 次
- Inference-Time Text-to-Video Alignment with Diffusion Latent Beam SearchYuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki FurutaNeurIPS 2025 · 被引用 50 次
- Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget ForcingJaihoon Kim, Taehoon Yoon, Jisung Hwang, Minhyuk SungNeurIPS 2025 · 被引用 43 次
它引用的顶会 Paper48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- Diffusion Alignment as Variational Expectation-MaximizationJaewoo Lee, Minsu Kim, Sanghyeok Choi, Inhyuck Song 等ICLR 2026 · 被引用 2 次
- Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding OptimisationTaehoon Kim, Henry Gouk, Timothy M. HospedalesCVPR 2026 · 被引用 1 次
- Test-Time Guidance for Flow-Based Generative Models via Parallel Tempering on Source DistributionsShih-Hsin Wang, Joel Keller, Taos Transue, Drake Brown 等ICML 2026
- Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-FunctionHyeongyu Kang, Jaewoo Lee, Woocheol Shin, Kiyoung Om 等ICLR 2026 · 被引用 5 次
- Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion modelsVineet Jain, Kusha Sareen, Mohammad Pedramfar, Siamak RavanbakhshNeurIPS 2025 · 被引用 41 次
