DisenStudio: Customized Multi-Subject Text-to-Video Generation with Disentangled Spatial Control
Hong Chen, Xin Wang, Yipeng Zhang, Yuwei Zhou, Zeyang Zhang, Siao Tang, Wenwu Zhu
Abstract
Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attributebinding problems when the video is expected to contain multiple subjects. Furthermore, existing models struggle to assign the desired actions to the corresponding subjects (action-binding problem), failing to achieve satisfactory multi-subject generation performance. To tackle the problems, in this paper, we propose DisenStudio, a novel framework that can generate text-guided videos for customized multiple subjects, given few images for each subject. Specifically, DisenStudio enhances a pretrained diffusion-based text-to-video model with our proposed spatial-disentangled cross-attention mechanism to associate each subject with the desired action. Then the model is customized for the multiple subjects with the proposed motion-preserved disentangled finetuning, which involves three tuning strategies: multi-subject co-occurrence tuning, masked single-subject tuning, and multi-subject motion-preserved tuning. The first two strategies guarantee the subject occurrence and preserve their visual attributes, and the third strategy helps the model maintain the temporal motion-generation ability when finetuning on static images. We conduct extensive experiments to demonstrate our proposed DisenStudio significantly outperforms existing methods in various metrics. Additionally, we show that DisenStudio can be used as a powerful tool for various controllable generation applications. 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5956d8f7-f4a7-44a2-94de-8d5056d6eb70Cited by top-tier papers13
- LLM4DyG: Can Large Language Models Solve Spatial-Temporal Problems on Dynamic Graphs?Zeyang Zhang, Xin Wang, Ziwei Zhang, Haoyang Li et al.KDD 2024 · 32 citations
- BindWeave: Subject-Consistent Video Generation via Cross-Modal IntegrationZhaoyang Li, Dongjun Qian, Kai Su, qishuai diao et al.ICLR 2026 · 23 citations
- PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and EnhancementTeng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang et al.NeurIPS 2025 · 15 citations
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLMZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen et al.AAAI 2025 · 6 citations
- DreamRelation: Relation-Centric Video CustomizationYujie Wei, Shiwei Zhang, Hangjie Yuan, Biao Gong et al.ICCV 2025 · 5 citations
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image GenerationHong Chen, Yipeng Zhang, Simin Wu, Xin Wang et al.ICLR 2024 · 81 citations
- VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion ModelsChi-Pin Huang, Yen-Siang Wu, Hung-Kai Chung, Kai-Po Chang et al.CVPR 2025
- CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition AbilitiesTao Wu, Yong Zhang, Xintao Wang, Xianpan Zhou et al.AAAI 2025 · 62 citations
- DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image GenerationJing He, Haodong Li, Yongzhe Hu, Guibao Shen et al.ICLR 2025
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang et al.CVPR 2026 · 9 citations
