Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
Feng Liang, Haoyu Ma, Zecheng He, Tingbo Hou, Ji Hou, Kunpeng Li, Xiaoliang Dai, Felix Juefei-Xu, Samaneh Azadi, Animesh Sinha, Peizhao Zhang, Peter Vajda, Diana Marculescu
Abstract
Video personalization, which generates customized videos using reference images, has gained significant attention. However, prior methods typically focus on singleconcept personalization, limiting broader applications that require multi-concept integration. Attempts to extend these models to multiple concepts often lead to identity blending, which results in composite characters with fused attributes from multiple sources. This challenge arises due to the lack of a mechanism to link each concept with its specific reference image. We address this with anchored prompts, which embed image anchors as unique tokens within text prompts, guiding accurate referencing during generation. Additionally, we introduce concept embeddings to encode the order of reference images. Our approach, Movie Weaver, seamlessly weaves multiple concepts-including face, body, and animal images-into one video, allowing flexible combinations in a single model. The evaluation shows that Movie Weaver outperforms existing methods for multi-concept video personalization in identity preservation and overall quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 956ee07d-e3ab-4212-8ee3-c2990f2f487eCited by top-tier papers7
- Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentLijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen et al.ICCV 2025 · 128 citations
- Phantom-Data: Towards a General Subject-Consistent Video Generation DatasetZhuowei Chen, Bingchuan Li, Tianxiang Ma, Lijie Liu et al.ICLR 2026 · 20 citations
- MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video GenerationRun Ling, Ke Cao, Jian Lu, Ao Ma et al.AAAI 2026 · 4 citations
- MoCha: Towards Movie-Grade Talking Character GenerationCong Wei, Bo Sun, Haoyu Ma, Ji Hou et al.NeurIPS 2025 · 2 citations
- HOComp: Interaction-Aware Human-Object CompositionDong Liang, Jinyuan Jia, Yuhao Liu, Rynson W. H. LauNeurIPS 2025 · 1 citation
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsGihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee et al.CVPR 2024
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject DisentanglementYufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang et al.ICLR 2026 · 20 citations
- Multi-subject Open-set Personalization in Video GenerationTsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang et al.CVPR 2025
- AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual ReferencesJiahao Wang, Hualian Sheng, Sijia Cai, Yuxiao Yang et al.CVPR 2026 · 1 citation
- MC^2: Multi-concept Guidance for Customized Multi-concept GenerationJiaxiu Jiang, Yabo Zhang, Kailai Feng, Xiaohe Wu et al.CVPR 2025
