EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
Rang Meng, Xingyu Zhang, Yuming Li, Chenguang Ma
Abstract
Recent work on human animation usually involves audio, pose, or movement maps conditions, thereby achieves vivid animation quality. However, these methods often face practical challenges due to extra control conditions, cumbersome condition injection modules, or limitation to head region driving. Hence, we ask if it is possible to achieve striking half-body human animation while simplifying unnecessary conditions. To this end, we propose a half-body human animation method, dubbed EchoMimicV2, that leverages a novel Audio-Pose Dynamic Harmonization strategy, including Pose Sampling and Audio Diffusion, to enhance half-body details, facial and gestural expressiveness, and meanwhile reduce conditions redundancy. To compensate for the scarcity of half-body data, we utilize Head Partial Attention to seamlessly accommodate headshot data into our training framework, which can be omitted during inference, providing a free lunch for animation. Furthermore, we design the Phase-specific Denoising Loss to guide motion, detail, and low-level quality for animation in specific phases, respectively. Besides, we also present a novel benchmark for evaluating the effectiveness of half-body human animation. Extensive experiments and analyses demonstrate that EchoMimicV2 surpasses existing methods in both quantitative and qualitative evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Let Them Talk: Audio-Driven Multi-Person Conversational Video GenerationZhe Kong, Feng Gao, Yong Zhang, Zhuoliang Kang et al.NeurIPS 2025 · 73 citations
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng et al.CVPR 2026 · 40 citations
- Instilling an Active Mind in Avatars via Cognitive SimulationJianwen Jiang, Weihong Zeng, Zerong Zheng, Jiaqi Yang et al.ICLR 2026 · 26 citations
- StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human AvatarsZhiyao Sun, Ziqiao Peng, Yifeng Ma, Yi Chen et al.CVPR 2026 · 26 citations
- EchoMimicV3: 1.3B Parameters Are All You Need for Unified Multi-Modal and Multi-Task Human AnimationRang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng et al.AAAI 2026 · 24 citations
Builds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
Related papers
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsZhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li et al.AAAI 2025 · 197 citations
- Bringing Your Portrait to 3D PresenceJiawei Zhang, Lei Chu, Jiahao Li, Zhenyu Zang et al.CVPR 2026 · 3 citations
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsYuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu et al.CVPR 2026 · 3 citations
- CyberHost: A One-stage Diffusion Framework for Audio-driven Talking Body GenerationGaojie Lin, Jianwen Jiang, Chao Liang, Tianyun Zhong et al.ICLR 2025
- ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion ModelingQuanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu et al.NeurIPS 2024 · 21 citations
