Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
Bowen Xue, Zheng-Peng Duan, Qixin Yan, Wenjing Wang, Hao Liu, Chun-Le Guo, Chongyi Li, Chen Li, Jing Lyu
Abstract
Generating high-fidelity human videos that match user-specified identities is important yet challenging in the field of generative AI. Existing methods often rely on an excessive number of training parameters and lack compatibility with other AIGC tools. In this paper, we propose Stand-In, a lightweight and plug-and-play framework for identity preservation in video generation. Specifically, we introduce a conditional image branch into the pre-trained video generation model. Identity control is achieved through restricted self-attentions with conditional position mapping. Thanks to these designs, which greatly preserve the pre-trained prior of the video generation model, our approach is able to outperform other full-parameter training methods in video quality and identity preservation, even with just 1% additional parameters and only 2000 training pairs. Moreover, our framework can be seamlessly integrated for other tasks, such as subject-driven video generation, pose-referenced video generation, stylization, and face swapping.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7753749-4954-4f09-96b9-da78f496527eCited by top-tier papers8
- Scaling Zero-Shot Reference-to-Video GenerationZijian Zhou, Shikun Liu, Haozhe Liu, Haonan Qiu et al.CVPR 2026 · 10 citations
- EgoX: Egocentric Video Generation from a Single Exocentric VideoTaewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park et al.CVPR 2026 · 9 citations
- Lynx: Towards High-Fidelity Personalized Video GenerationShen Sang, Tiancheng Zhi, Tianpei Gu, Jing Liu et al.CVPR 2026 · 9 citations
- Identity-Preserving Image-to-Video Generation via Reward-Guided OptimizationLiao Shen, Wentao Jiang, Yiran Zhu, Jiahe Li et al.CVPR 2026 · 8 citations
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video GenerationPanwang Pan, Jingjing Zhao, Yuchen Lin, Chenguo Lin et al.CVPR 2026 · 5 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual ReferencesJiahao Wang, Hualian Sheng, Sijia Cai, Yuxiao Yang et al.CVPR 2026 · 1 citation
- A Latent Transformer for Disentangled Face Editing in Images and VideosXu Yao, Alasdair Newson, Yann Gousseau, Pierre HellierICCV 2021 · 97 citations
- I2V-Adapter: A General Image-to-Video Adapter for Diffusion ModelsXun Guo, Mingwu Zheng, Liang Hou, Yuan Gao et al.SIGGRAPH 2024 · 26 citations
- PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic DegradationHengjia Li, Haonan Qiu, Shiwei Zhang, Xiang Wang et al.ICCV 2025 · 1 citation
- MagicMirror: ID-Preserved Video Generation in Video Diffusion TransformersYuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng et al.ICCV 2025 · 1 citation
