BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
Zhaoyang Li, Dongjun Qian, Kai Su, qishuai diao, Xiangyang Xia, Chang Liu, Wenfei Yang, Tianzhu Zhang, Zehuan Yuan
摘要
Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in subject-consistent video generation due to an inherent difficulty in parsing prompts that specify complex spatial relationships, temporal logic, and interactions among multiple subjects. To address this issue, we propose BindWeave, a unified framework that handles a broad range of subject-to-video scenarios from single-subject cases to complex multi-subject scenes with heterogeneous entities. To bind complex prompt semantics to concrete visual subjects, we introduce an MLLM-DiT framework in which a pretrained multimodal large language model performs deep cross-modal reasoning to ground entities and disentangle roles, attributes, and interactions, yielding subject-aware hidden states that condition the diffusion transformer for high-fidelity subject-consistent video generation. Experiments on the OpenS2V benchmark demonstrate that our method achieves superior performance across subject consistency, naturalness, and text relevance in generated videos, outperforming existing open-source and commercial models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video GenerationXu Guo, Fulong Ye, Qichao Sun, Liyang Chen 等ICML 2026 · 被引用 16 次
- Scaling Zero-Shot Reference-to-Video GenerationZijian Zhou, Shikun Liu, Haozhe Liu, Haonan Qiu 等CVPR 2026 · 被引用 10 次
- AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video GenerationSharath Girish, Viacheslav Ivanov, Tsai-Shien Chen, Hao Chen 等CVPR 2026 · 被引用 4 次
- MiVE: Multiscale Vision-language features for reference-guided video EditingTong Wang, Meng Zou, WU CHENGJING, Xiaochao Qu 等ICML 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- UniVideo: Unified Understanding, Generation, and Editing for VideosCong Wei, Quande Liu, Zixuan Ye, Qiulin Wang 等ICLR 2026 · 被引用 90 次
- Composing Concepts from Images and Videos via Concept-prompt BindingXianghao Kong, Zeyu Zhang, Yuwei Guo, Zhuoran Zhao 等CVPR 2026 · 被引用 2 次
- LumosX: Relate Any Identities with Their Attributes for Personalized Video GenerationJiazheng Xing, Fei Du, Hangjie Yuan, Pengwei Liu 等ICLR 2026 · 被引用 6 次
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLMZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen 等AAAI 2025 · 被引用 6 次
- Compositional 3D-aware Video Generation with LLM DirectorHanxin Zhu, Tianyu He, Anni Tang, Junliang Guo 等NeurIPS 2024 · 被引用 19 次
