Circular-DPO: Aligning Multi-Stage 3D Generative Models via Preference Feedback Loop
Zejian Li, Jiarui Ma, Han Xu, Weiting Zheng, Yangrui Zhu, Chenye Meng, Pei Chen, Ling Yang, Zhiyuan Yang, Changyuan Yang, Guang Yang, Immanuel Koh, Lingyun Sun
摘要
Multi-stage generative models have shown great promise in 3D content creation due to focused generation of structure or texture in different stages, but their outputs often fail to align with human preferences. The key bottleneck to apply alignment methods is the presence of non-differentiable operations between generative stages.This disconnection stops preference signals applied to the final output from being backpropagated to the crucial, early stages of generation, while simple separated stage-wise alignment leads to texture-geometry inconsistency.To address this challenge, we introduce Circular-DPO, which builds a preference feedback loop to align multi-stage 3D generation models to human preference.Our method first applies Direct Preference Optimization (DPO) to refine the final 3D asset.We then construct new preference pairs by sampling and decoding the assets generated by the optimized model.These newly-formed pairs are used to train the preceding generative stage, effectively creating a feedback loop that bridges the non-differentiable gap. Furthermore, to enhance robustness against noisy data, we introduce a quality-aware weighting mechanism that prioritizes reliable preference pairs during training. Experiments demonstrate that our approach improves the alignment of generated 3D content with human preferences by enabling holistic, multi-stage optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- Rethinking DPO-Style Diffusion Aligning FrameworksXun Wu, Shaohan Huang, Lingjie Jiang, Furu WeiICCV 2025 · 被引用 4 次
- DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference OptimizationZhenglin Zhou, Xiaobo Xia, Fan Ma, Hehe Fan 等ICML 2025
- Curriculum Direct Preference Optimization for Diffusion and Consistency ModelsFlorinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe 等CVPR 2025
- DSO: Aligning 3D Generators with Simulation Feedback for Physical SoundnessRuining Li, Chuanxia Zheng, Christian Rupprecht, Andrea VedaldiICCV 2025 · 被引用 2 次
- A Gradient Guidance Perspective on Stepwise Preference Optimization for Diffusion ModelsJoshua Tian Jin Tee, Hee Suk Yoon, Abu Hanif Muhammad Syarubany, Eunseop Yoon 等NeurIPS 2025 · 被引用 1 次
