OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
Yuanhao Cai, He Zhang, Xi Chen, Jinbo Xing, Yiwei Hu, Yuqian Zhou, Kai Zhang, Zhifei Zhang, Soo Ye Kim, Tianyu Wang, Yulun Zhang, Xiaokang Yang
摘要
Existing feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as depth, mask, camera, and text prompts to control and edit the subject in the customized video is still less explored. In this paper, we first propose a data construction pipeline, VideoCus-Factory, to produce training data pairs for multi-subject customization from raw videos without labels and control signals such as depth-to-video and mask-to-video pairs. Based on our constructed data, we develop an Image-Video Transfer Mixed (IVTM) training with image editing data to enable instructive editing for the subject in the customized video. Then we propose a diffusion Transformer framework, OmniVCus, with two embedding mechanisms, Lottery Embedding (LE) and Temporally Aligned Embedding (TAE). LE enables inference with more subjects by using the training subjects to activate more frame embeddings. TAE encourages the generation process to extract guidance from temporally aligned control signals by assigning the same frame embeddings to the control and noise tokens. Experiments demonstrate that our method significantly surpasses state-of-the-art methods in both quantitative and qualitative evaluations. Video demos are at our project page: https://caiyuanhao1998.github.io/project/OmniVCus/. Our code, models, data are released at https://github.com/caiyuanhao1998/Open-OmniVCus
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang 等ICLR 2026 · 被引用 56 次
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video GenerationJiehui Huang, Yuechen Zhang, Xu He, Yuan Gao 等CVPR 2026 · 被引用 12 次
- VideoCoF: Unified Video Editing with Temporal ReasonerXiangpeng Yang, Ji Xie, Yiyuan Yang, Yue Ma 等CVPR 2026 · 被引用 6 次
- MATRIX: Mask Track Alignment for Interaction-aware Video GenerationSiyoon Jin, Seongchan Kim, Jae Ho Lee, Dahyun Chung 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video GenerationXu Guo, Fulong Ye, Qichao Sun, Liyang Chen 等ICML 2026 · 被引用 16 次
- OmniVDiff: Omni Controllable Video Diffusion for Generation and UnderstandingDianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu 等AAAI 2026 · 被引用 14 次
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao 等ICML 2026 · 被引用 9 次
- Less-to-More Generalization: Unlocking More Controllability by In-Context GenerationShaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng 等ICCV 2025 · 被引用 11 次
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual GenerationJunke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 132 次
