Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset
Zhuowei Chen, Bingchuan Li, Tianxiang Ma, Lijie Liu, Mingcong Liu, Yunsheng Jiang, Gen Li, Xinghui Li, Liyang Chen, Siyu Zhou, Qian He, Xinglong Wu
Abstract
Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructions. This limitation, commonly known as the copy-paste problem, arises from the widely used in-pair training paradigm. This approach inherently entangles subject identity with background and contextual attributes by sampling reference images from the same scene as the target video. To address this issue, we introduce Phantom-Data, the first general-purpose cross-pair subject-to-video consistency dataset, containing approximately one million identity-consistent pairs across diverse categories. Our dataset is constructed via a three-stage pipeline: (1) a general and input-aligned subject detection module, (2) large-scale cross-context subject retrieval from more than 53 million videos and 3 billion images, and (3) prior-guided identity verification to ensure visual consistency under contextual variation. Comprehensive experiments show that training with Phantom-Data significantly improves prompt alignment and visual quality while preserving identity consistency on par with in-pair baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6279afc-0232-4c70-9b6d-6433d83848daCited by top-tier papers5
- DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video GenerationXu Guo, Fulong Ye, Qichao Sun, Liyang Chen et al.ICML 2026 · 16 citations
- Scaling Zero-Shot Reference-to-Video GenerationZijian Zhou, Shikun Liu, Haozhe Liu, Haonan Qiu et al.CVPR 2026 · 10 citations
- Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video GenerationBinyuan Huang, Yuning Lu, Weinan Jia, Hualiang Wang et al.CVPR 2026 · 3 citations
- MV-S2V: Multi-View Subject-Consistent Video GenerationZiyang Song, Xinyu Gong, Bangya Liu, Zelin ZhaoSIGGRAPH 2026 · 1 citation
- Human-Centric Video Generation via Collaborative Multi-Modal ConditioningLiyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li et al.AAAI 2026 · 1 citation
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama et al.ICML 2024 · 464 citations
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan et al.NeurIPS 2025 · 284 citations
Related papers
- Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentLijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen et al.ICCV 2025 · 128 citations
- ReactID: Synchronizing Realistic Actions and Identity in Personalized Video GenerationWei Li, Yiheng Zhang, Fuchen Long, Zhaofan Qiu et al.ICLR 2026
- ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video GenerationMingyang Wu, Ashirbad Mishra, Soumik Dey, Shuo Xing et al.CVPR 2026 · 8 citations
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject DisentanglementYufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang et al.ICLR 2026 · 20 citations
- CI-VID: A Coherent Interleaved Text-Video DatasetYiming Ju, Jijin Hu, Zhengxiong Luo, Haoge Deng et al.CVPR 2026 · 3 citations
