Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
Harold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang, Ser Nam Lim
摘要
Recent advancements in video generation have enabled the creation of high-quality, visually compelling videos. However, generating videos that adhere to the laws of physics remains a critical challenge for applications requiring realism and accuracy. In this work, we propose PhysHPO, a novel framework for Hierarchical Cross-Modal Direct Preference Optimization, to tackle this challenge by enabling finegrained preference alignment for physically plausible video generation. PhysHPO optimizes video alignment across four hierarchical granularities: a) Instance Level, aligning the overall video content with the input prompt; b) State Level, ensuring temporal consistency using boundary frames as anchors; c) Motion Level, modeling motion trajectories for realistic dynamics; and d) Semantic Level, maintaining logical consistency between narrative and visuals. Recognizing that real-world videos are the best reflections of physical phenomena, we further introduce an automated data selection pipeline to efficiently identify and utilize "good data" from existing large-scale text-video datasets, thereby eliminating the need for costly and time-intensive dataset construction. Extensive experiments on both physicsfocused and general capability benchmarks demonstrate that PhysHPO significantly improves physical plausibility and overall video generation quality of advanced models. To the best of our knowledge, this is the first work to explore fine-grained preference alignment and data selection for video generation, paving the way for more realistic and human-preferred video generation paradigms. PhysHPO
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ProPhy: Progressive Physical Alignment for Dynamic World SimulationZijun Wang, Panwen Hu, Jing Wang, Terry Jingchen Zhang 等CVPR 2026 · 被引用 14 次
- Chain of Event-Centric Causal Thought for Physically Plausible Video GenerationZixuan Wang, Yixin Hu, Haolan Wang, Feng Chen 等CVPR 2026 · 被引用 8 次
- Show, Don't Tell: Morphing Latent Reasoning into Image GenerationHarold Haodong Chen, Xinxiang Yin, Wenjie Shu, Hongfei (Faye) Zhang 等ICML 2026 · 被引用 7 次
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVerPukun Zhao, Longxiang Wang, Miaowei Wang, Chen Chen 等AAAI 2026 · 被引用 2 次
- Find, Fix, Reason: Context Repair for Video ReasoningHaojian Huang, Chuanyu Qin, Yinchuan Li, YINGCONG CHENICML 2026 · 被引用 1 次
它引用的顶会 Paper49
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
相关 Paper
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video ModelsHaojian Huang, Haodong Chen, Shengqiong Wu, Meng Luo 等ICML 2025
- PHANTOM: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical DynamicsYing Shen, Jerry Xiong, Tianjiao Yu, Ismini LourentzouCVPR 2026 · 被引用 12 次
- VideoDPO: Omni-Preference Alignment for Video Diffusion GenerationRuntao Liu, Haoyu Wu, Ziqiang Zheng, Chen Wei 等CVPR 2025
- VPO: Aligning Text-to-Video Generation Models with Prompt OptimizationJiale Cheng, Ruiliang Lyu, Xiaotao Gu, Xiao Liu 等ICCV 2025 · 被引用 3 次
- PhysVid: Physics Aware Local Conditioning for Generative Video ModelsSaurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram ZonoozCVPR 2026 · 被引用 6 次
