Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
Binyuan Huang, Yuning Lu, Weinan Jia, Hualiang Wang, Mu Liu, Daiqing Yang
Abstract
Recent proprietary models such as Sora2 demonstrate promising progress in generating multi-shot videos conditioned on multiple reference characters. However, academic research on this problem remains limited. We study this task and identify a core challenge: when reference images exhibit highly similar appearances, the model often suffers from reference confusion, where semantically similar tokens degrade the model's ability to retrieve the correct context. To address this, we introduce PoCo (Position Embedding as a Context Controller), which incorporates position encoding as additional context control beyond semantic retrieval. By employing side information of tokens, PoCo enables precise token-level matching while preserving implicit semantic consistency modeling. Building on PoCo, we develop a multi-reference and multi-shot video generation model capable of reliably controlling characters with extremely similar visual traits. Extensive experiments demonstrate that PoCo improves cross-shot consistency and reference fidelity compared with various baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6271487f-92df-4feb-9ad1-7a6af478d1eeBuilds on19
- Composer: Creative and Controllable Image Synthesis with Composable ConditionsLianghua Huang, Di Chen, Yu Liu, Yujun Shen et al.ICML 2023 · 371 citations
- Cones: Concept Neurons in Diffusion Models for Customized GenerationZhiheng Liu, Ruili Feng, Kai Zhu, Yifei Zhang et al.ICML 2023 · 164 citations
- Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentLijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen et al.ICCV 2025 · 128 citations
- Video-P2P: Video Editing with Cross-Attention ControlShaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin et al.CVPR 2024 · 99 citations
- Mixture of Contexts for Long Video GenerationShengqu Cai, Ceyuan Yang, Lvmin Zhang, Yuwei Guo et al.ICLR 2026 · 92 citations
Related papers
- Gloria: Consistent Character Video Generation via Content AnchorsYuhang Yang, Fan Zhang, Huaijin Pi, Ailing Zeng et al.CVPR 2026 · 3 citations
- MultiShotMaster: A Controllable Multi-Shot Video Generation FrameworkQinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian et al.CVPR 2026 · 33 citations
- ReRoPE: Repurposing RoPE for Relative Camera ControlChunyang Li, Yuanbo Yang, Jiahao Shao, Hongyu Zhou et al.SIGGRAPH 2026 · 2 citations
- MV-S2V: Multi-View Subject-Consistent Video GenerationZiyang Song, Xinyu Gong, Bangya Liu, Zelin ZhaoSIGGRAPH 2026 · 1 citation
- Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored PromptsFeng Liang, Haoyu Ma, Zecheng He, Tingbo Hou et al.CVPR 2025
