Real-Time Motion-Controllable Autoregressive Video Diffusion
Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang
摘要
Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first finetune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the Markov property through a Self-Rollout mechanism and accelerates training by selectively introducing stochasticity in denoising steps. Extensive experiments demonstrate that AR-Drag achieves high visual fidelity and precise motion alignment, significantly reducing latency compared with state-of-the-art motion-controllable VDMs, while using only 1.3B parameters. Additional visualizations can be found on our project page: https://kesenzhao.github. io/AR-Drag.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng 等ICLR 2026 · 被引用 130 次
- Adaptive Stochastic Coefficients for Accelerating Diffusion SamplingRuoyu Wang, Beier Zhu, Junzhi Li, Liangyu Yuan 等NeurIPS 2025 · 被引用 8 次
- DragNeXt: Rethinking Drag-Based Image EditingYuan Zhou, Junbao Zhou, Qingshan Xu, Kesen Zhao 等AAAI 2026 · 被引用 7 次
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!Junbao Zhou, Yuan Zhou, Kesen Zhao, Qingshan Xu 等ICLR 2026 · 被引用 7 次
- Free Lunch for Stabilizing Rectified Flow InversionChenru Wang, Beier Zhu, Chi ZhangICLR 2026 · 被引用 6 次
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz 等NeurIPS 2024 · 被引用 751 次
相关 Paper
- DragEntity: Trajectory Guided Video Generation using Entity and Positional RelationshipsZhang Wan, Sheng Tang, Jiawei Wei, Ruize Zhang 等ACM MM 2024 · 被引用 4 次
- MV-Diffusion: Motion-aware Video Diffusion ModelZijun Deng, Xiangteng He, Yuxin Peng, Xiongwei Zhu 等ACM MM 2023 · 被引用 20 次
- DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion ControlKaifeng Zhao, Gen Li, Siyu TangICLR 2025 · 被引用 1 次
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 等NeurIPS 2025 · 被引用 628 次
- AR-Diffusion: Asynchronous Video Generation with Auto-Regressive DiffusionMingzhen Sun, Weining Wang, Gen Li, Jiawei Liu 等CVPR 2025
