PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
Minh-Quan Le, Gaurav Mittal, Cheng Zhao, Xianfeng GU, Samaras Dimitris, Mei Chen
Abstract
Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality and semantic alignment of generated videos. However, recent methods either rely on large-scale human preference annotations or operate on misaligned embeddings from pre-trained vision-language models, leading to limited scalability or suboptimal supervision. We present PISCES, an annotation-free post-training algorithm that addresses these limitations via a novel Dual Optimal Transport (OT)-aligned Rewards module. To align reward signals with human judgment, PISCES uses OT to bridge text and video embeddings at both distributional and discrete token levels, enabling reward supervision to fulfill two objectives: (i) a Distributional OT-aligned Quality Reward that captures overall visual quality and temporal coherence; and (ii) a Discrete Token-level OT-aligned Semantic Reward that enforces semantic, spatio-temporal correspondence between text and video tokens. To our knowledge, PISCES is the first to improve annotation-free reward supervision in generative post-training through the lens of OT. Experiments on both short- and long-video generation show that PISCES outperforms both annotation-based and annotation-free methods on VBench across Quality and Semantic scores, with human preference studies further validating its effectiveness. We show that the Dual OT-aligned Rewards module is compatible with multiple optimization paradigms, including direct backpropagation and reinforcement learning fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a37bfba-4a3e-4765-a116-b8c6b7a598b6Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement LearningMinh-Tung Luu, Hwanhee Kim, Younghwan Lee, Chang D. YooICML 2026 · 1 citation
- Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward ModelYuan Wang, Borui Liao, Huijuan Huang, Jinda Lu et al.CVPR 2026 · 5 citations
- Inference-Time Text-to-Video Alignment with Diffusion Latent Beam SearchYuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki FurutaNeurIPS 2025 · 50 citations
- Free2 Guide: Training-Free Text-to-Video Alignment Using Image LVLMJaemin Kim, Bryan Sangwoo Kim, Jong Chul YeICCV 2025 · 1 citation
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward ModelsJiesong Lian, Ruizhe Zhong, Zixiang Zhou, Xiaoyue Mi et al.CVPR 2026 · 3 citations
