TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment
Jin Wang, Jianxiang Lu, Guangzheng Xu, Comi Chen, Haoyu Yang, zhenzhen qin, Peng Chen, Mingtao Chen, Zhichao hu, Longhuang Wu, Shuai Shao, Qinglin Lu, Ping Luo
Abstract
Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video generation. However, we find that directly applying these techniques to image-to-video (I2V) models often fails to yield consistent reward improvements. To address this limitation, we present TAGRPO, a robust post-training framework for I2V models inspired by contrastive learning. Our approach is grounded in the observation that rollout videos generated from identical initial noise provide superior guidance for optimization. Leveraging this insight, we propose a novel GRPO loss applied to intermediate latents, encouraging direct alignment with high-reward trajectories while maximizing distance from low-reward counterparts. Furthermore, we introduce a memory bank for rollout videos to enhance diversity and reduce computational overhead. Despite its simplicity, TAGRPO achieves significant improvements over DanceGRPO in I2V generation. The deliverables will be updated here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- Principled RL for Flow Matching Emerges from the Chunk-level Policy OptimizationYifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du et al.ICML 2026 · 12 citations
- Seeing What Matters: Visual Preference Policy Optimization for Visual GenerationZiqi Ni, Yuanzhi Liang, Rui Li, Yi Zhou et al.CVPR 2026 · 9 citations
- Diverse Video Generation with Determinantal Point Process-Guided Policy OptimizationTahira Kazimi, Connor Dunlop, Pinar YanardagCVPR 2026 · 4 citations
- Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow ModelsDailan He, Guanlin Feng, Xingtong Ge, Yazhe Niu et al.CVPR 2026 · 15 citations
- Learning What to Trust: Bayesian Prior-Guided Optimization for Visual GenerationRuiying Liu, Yuanzhi Liang, Haibin Huang, Tianshu Yu et al.CVPR 2026 · 4 citations
