PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation
Hengjia Li, Haonan Qiu, Shiwei Zhang, Xiang Wang, Yujie Wei, Zekun Li, Yingya Zhang, Boxi Wu, Deng Cai
Abstract
The current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies in maintaining high ID fidelity consistently while preserving the original motion dynamic and semantic following after the identity injection. Current video identity customization methods mainly rely on reconstructing given identity images on text-to-image models, which have a divergent distribution with the T2V model. This process introduces a tuning-inference gap, leading to dynamic and semantic degradation. To tackle this problem, we propose a novel framework, dubbed , that applies a mixture of reward supervision on synthesized videos instead of the simple reconstruction objective on images. Specifically, we first incorporate identity consistency reward to effectively inject the reference's identity without the tuning-inference gap. Then we propose a novel semantic consistency reward to align the semantic distribution of the generated videos with the original T2V model, which preserves its dynamic and semantic following capability during the identity injection. With the non-reconstructive reward training, we further employ simulated prompt augmentation to reduce overfitting by supervising generated results in more semantic scenarios, gaining good robustness even with only a single reference image. Extensive experiments demonstrate our method's superiority in delivering high identity faithfulness while preserving the inherent video generation qualities of the original T2V model, outshining prior methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4df398c2-a895-4080-958d-8db3f052f4cfCited by top-tier papers11
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao et al.NeurIPS 2025 · 25 citations
- SMRABooth: Subject and Motion Representation Alignment for Customized Video GenerationXuancheng Xu, Yaning Li, Sisi You, Bing-Kun BaoCVPR 2026 · 11 citations
- Identity-Preserving Image-to-Video Generation via Reward-Guided OptimizationLiao Shen, Wentao Jiang, Yiran Zhu, Jiahe Li et al.CVPR 2026 · 8 citations
- InfiniDreamer: Arbitrarily Long Human Motion Generation Via Segment Score DistillationWenjie Zhuo, Fan Ma, Hehe FanICCV 2025 · 6 citations
- RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlTeng Li, Guangcong Zheng, Rui Jiang, Shuigen Zhan et al.ICCV 2025 · 5 citations
Builds on20
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan et al.ICCV 2023 · 770 citations
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
Related papers
- Magicid: Hybrid Preference Optimization for Id-Consistent and Dynamic-Preserved Video CustomizationHengjia Li, Lifan Jiang, Xi Xiao, Tianyang Wang et al.ICCV 2025 · 2 citations
- I2V-Adapter: A General Image-to-Video Adapter for Diffusion ModelsXun Guo, Mingwu Zheng, Liang Hou, Yuan Gao et al.SIGGRAPH 2024 · 26 citations
- Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentLijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen et al.ICCV 2025 · 128 citations
- AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual ReferencesJiahao Wang, Hualian Sheng, Sijia Cai, Yuxiao Yang et al.CVPR 2026 · 1 citation
- DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video CustomizationWenchuan Wang, Mengqi Huang, Yijing Tu, Zhendong MaoICCV 2025 · 3 citations
