MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation
Xirui Hu, Yanbo Ding, Jiahao Wang, Tingting Shi, Yali Wang, Guo Zhi Zhi, Weizhan Zhang
Abstract
a conference paper Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Existing methods struggle to generalize to multi-humanoid scenarios, which involve diverse humanoid forms, complex interactions, and frequent occlusions. We address this gap with two key innovations. First, we introduce unified motion representations that extract identity-agnostic motions and explicitly bind them to corresponding characters, enabling generalization across diverse humanoid forms and seamless extension to multi-humanoid scenarios. Second, we propose a holistic 4D-anchored paradigm that constructs a shared 4D space to fuse motion representations with video latents, and further reinforces this process with hierarchical 4D-level supervision to better handle interactions and occlusions. We instantiate these ideas in MotionWeaver, an end-to-end framework for multi-humanoid image animation. To support this setting, we curate a 46-hour dataset of multi-human videos with rich interactions, and construct a 300-video benchmark featuring paired humanoid characters. Quantitative and qualitative experiments demonstrate that MotionWeaver not only achieves state-of-the-art results on our benchmark but also generalizes effectively across diverse humanoid forms, complex interactions, and challenging multi-humanoid scenarios. a conference paper pervision (H4S): couples occlusion supervision at high-noise steps with motion-level supervision at low-noise steps, providing 4D motion supervision and mitigating overfitting to human appearance. Overall, the contribution of MotionWeaver lies in introducing novel unified motion representations and formulating a holistic 4D-anchored paradigm where motion extraction, motion-latent fusion, and training supervision are consistently grounded in 4D space. Building on these designs, MotionWeaver constitutes an end-to-end framework that supports multi-humanoid animation and robustly addresses interactions and occlusions. To enhance the training of our model, we further curate a dataset containing 46 hours of multi-human videos, referred to as MultiHuman46, which features diverse interaction patterns and scenes. Additionally, we introduce DualDynamics, a benchmark of 300 videos, each showcasing two humanoid characters engaged in interaction-rich scenarios. These videos have undergone a rigorous filtering process to ensure quality. Quantitative and qualitative experiments demonstrate that MotionWeaver surpasses state-of-the-art methods, showcasing its generalization ability, identity preservation, and motion consistency in multi-humanoid scenarios. Our main contributions are summarized as follows: • We propose MotionWeaver, a novel framework built upon the unified motion representations and the 4D-anchored paradigm, designed for multi-humanoid image animation involving diverse humanoid forms, rich interactions, and frequent occlusions. • We introduce UCC to obtain unified motion representations, HSI and H4S to effectively construct a shared 4D space for fusing motion representations with video latents. • We curate the MultiHuman46 dataset, which encompasses 46 hours of multi-human videos, and create DualDynamics, a benchmark comprising 300 videos of multiple humanoid characters in interaction-rich scenarios. Extensive experiments demonstrate that MotionWeaver surpasses state-of-the-art methods in multi-humanoid scenarios. 2 RELATED WORK 2.1 DIFFUSION TRANSFORMERS Diffusion Transformers (DiTs) replace the traditional U-Net architecture with a Transformer model to denoise latent representations (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eaaaf95d-c353-4026-b1ce-657e35057606Cited by top-tier papers1
Ask how each one uses itBuilds on36
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang et al.AAAI 2024 · 318 citations
- DreamPose: Fashion Image-to-Video Synthesis via Stable DiffusionJohanna Suvi Karras, Aleksander Holynski, Ting-Chun Wang, Ira Kemelmacher-ShlizermanICCV 2023 · 224 citations
Related papers
- MTVCraft: Tokenizing 4D Motion for Arbitrary Character AnimationYanbo Ding, Xirui Hu, Guo Zhi, Yan Zhang et al.ICLR 2026 · 3 citations
- Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion ModelsMarc Benedí San Millán, Angela Dai, Matthias NießnerICLR 2026 · 3 citations
- MultiAnimate: Pose-Guided Image Animation Made ExtensibleYingcheng Hu, Haowen Gong, Chuanguang Yang, Zhulin An et al.CVPR 2026 · 6 citations
- DreamActor-M2: Universal Character Image Animation via Spatiotemporal In-Context LearningMingshuang Luo, Shuang Liang, Zhengkun Rong, Yuxuan Luo et al.SIGGRAPH 2026
- Motion 3-to-4: 3D Motion Reconstruction for 4D SynthesisHongyuan Chen, Xingyu Chen, Zexiang Xu, Anpei ChenCVPR 2026 · 17 citations
