Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
Liyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li, Zhuowei Chen, Lijie Liu, Xu He, Gen Li, Qian He, Zhiyong Wu
Abstract
Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, image, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of training data with paired triplet conditions and the difficulty of collaborating the sub-tasks of subject preservation and audio-visual sync with multimodal inputs. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct a high-quality dataset with diverse and paired text, reference images, and audio. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies. For the subject preservation task, to maintain the prompt following and visual generation abilities of the foundation model, we adopt the minimal-invasive image injection strategy. For the audio-visual sync task, besides the commonly adopted audio cross-attention layer, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multimodal inputs, building on previously acquired capabilities, we progressively incorporate the audio-visual sync task. During inference, for flexible and fine-grained multimodal control, we design a time-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- EgoX: Egocentric Video Generation from a Single Exocentric VideoTaewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park et al.CVPR 2026 · 9 citations
- From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative BootstrappingXu He, Haoxian Zhang, Hejia Chen, Changyuan Zheng et al.ICML 2026
- ExpPortrait: Expressive Portrait Generation via Personalized RepresentationJunyi Wang, Yudong Guo, Boyang Guo, Shengming Yang et al.CVPR 2026
- DiasR: Dual-Modal Identity-Anchored Sparse Routing for Efficient Multi-Subject Video GenerationYang-yang Li, Wu Liu, Jie Li, Xinchen Liu et al.ICML 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Improving Video Generation with Human FeedbackJie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan et al.NeurIPS 2025 · 284 citations
- Phantom: Subject-Consistent Video Generation via Cross-Modal AlignmentLijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen et al.ICCV 2025 · 128 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
Related papers
- OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video GenerationDonghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li et al.ICML 2026 · 3 citations
- MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio SynthesisHo Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya et al.CVPR 2025
- Harmony: Harmonizing Audio and Video Generation through Cross-Task SynergyTeng Hu, Zhentao Yu, Guozhen Zhang, Zihan Su et al.CVPR 2026 · 21 citations
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng et al.CVPR 2026 · 40 citations
- GENMO: A GENeralist Model for Human MOtionJiefeng Li, Jinkun Cao, Haotian Zhang, Davis Rempe et al.ICCV 2025 · 15 citations
