EvoID: Reinforced Evolution for Identity-Preserving Video Generation
Yiheng Zhang, Zhaofan Qiu, Zunxu Liu, Yingwei Pan, Ting Yao, Tao Mei
摘要
We present EvoID, a novel framework that reformulates Identity-Preserving Video Generation as a self-evolving process through Reinforcement Learning. Moving beyond the static paradigm of imitation learning, EvoID enables a generative model to actively learn and optimize the complex trade-offs between identity fidelity, motion naturalness, and temporal coherence. At the heart of our EvoID is a dynamic, dual-path reward mechanism, which acts as an intrinsic critic by adaptively combining objective metric indicators and MLLM-based holistic quality assessment. This allows the model to "evolve" its generation strategy, focusing on different aspects of quality at different stages of training. To ensure stable and coherent evolution, we anchor the exploring Student model with a frozen Teacher, preserving robust world priors while allowing for creative refinement when generating videos. Extensive experiments demonstrate the superiority of our proposal, and EvoID achieves the total score of 0.687 on the Human-Domain of OpenS2V-Eval dataset, surpassing 0.658 of the open-source VACE and 0.653 of the commercial Hailuo. Moreover, EvoID also obtains a new record of 0.718 on our newly minted MLLM-based metric, prioritizing human perception and more comprehensively reflecting video quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov 等ICLR 2024 · 被引用 816 次
相关 Paper
- AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual ReferencesJiahao Wang, Hualian Sheng, Sijia Cai, Yuxiao Yang 等CVPR 2026 · 被引用 1 次
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text InterpretationJiarui Wang, Huiyu Duan, Ziheng Jia, Zicheng Zhang 等ICML 2026 · 被引用 14 次
- RunawayEvil: Jailbreaking the Image-to-Video Generative Modelsyueming lyu, Rufan Qian, Yueming Lyu, Qinglong Liu 等CVPR 2026 · 被引用 7 次
- PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and EnhancementTeng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang 等NeurIPS 2025 · 被引用 15 次
- ReactID: Synchronizing Realistic Actions and Identity in Personalized Video GenerationWei Li, Yiheng Zhang, Fuchen Long, Zhaofan Qiu 等ICLR 2026
