DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
Ziyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Ashkan Mirzaei, Igor Gilitschenski, Sergey Tulyakov, Aliaksandr Siarohin
摘要
Direct Preference Optimization (DPO) has recently been applied as a post-training technique for text-to-video diffusion models. To obtain training data, annotators are asked to provide preferences between two videos generated from independent noise. However, this approach prohibits fine-grained comparisons, and we point out that it biases the annotators towards low-motion clips as they often contain fewer visual artifacts. In this work, we introduce DenseDPO, a method that addresses these shortcomings by making three contributions. First, we create each video pair for DPO by denoising corrupted copies of a ground truth video. This results in aligned pairs with similar motion structures while differing in local details, effectively neutralizing the motion bias. Second, we leverage the resulting temporal alignment to label preferences on short segments rather than entire clips, yielding a denser and more precise learning signal. With only one-third of the labeled data, DenseDPO greatly improves motion generation over vanilla DPO, while matching it in text alignment, visual quality, and temporal consistency. Finally, we show that DenseDPO unlocks automatic preference annotation using off-the-shelf Vision Language Models (VLMs): GPT accurately predicts segment-level preferences similar to task-specifically fine-tuned video reward models, and DenseDPO trained on these labels achieves performance close to using human labels. Additional results are available at https://snap-research.github.io/DenseDPO/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video GenerationKai Liu, Yanhao Zheng, Kai Wang, Shengqiong Wu 等ICLR 2026 · 被引用 24 次
- Towards Better Optimization For Listwise Preference in Diffusion ModelsJiamu Bai, Xin Yu, Meilong Xu, Weitao Lu 等ICLR 2026 · 被引用 8 次
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video GenerationPanwang Pan, Jingjing Zhao, Yuchen Lin, Chenguo Lin 等CVPR 2026 · 被引用 5 次
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait AnimationMengchao Wang, Qiang Wang, Fan Jiang, Mu XuAAAI 2026 · 被引用 4 次
- GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text RenderingXincheng Shuai, Ziye Li, Henghui Ding, Dacheng TaoCVPR 2026 · 被引用 4 次
它引用的顶会 Paper47
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion ModelsZitong Huang, Kaidong Zhang, Yukang Ding, Chao Gao 等CVPR 2026 · 被引用 3 次
- VideoDPO: Omni-Preference Alignment for Video Diffusion GenerationRuntao Liu, Haoyu Wu, Ziqiang Zheng, Chen Wei 等CVPR 2025
- Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference OptimizationTao Zhang, Cheng Da, Kun Ding, Huan Yang 等NeurIPS 2025 · 被引用 38 次
- InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model AlignmentYunhong Lu, Qichao Wang, Hengyuan Cao, Xierui Wang 等CVPR 2025
- Self-Supervised Direct Preference Optimization for Text-to-Image Diffusion ModelsLiang Peng, Boxi Wu, Haoran Cheng, Yibo Zhao 等NeurIPS 2025 · 被引用 2 次
