MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
Shurong Yang, Huadong Li, Juhao Wu, Minhao Jing, Linze Li, Renhe Ji, Jiajun Liang, Haoqiang Fan, Jin Wang
摘要
Diffusion models have demonstrated superior performance in portrait animation.
However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control.
This challenge arises from the difficulty in balancing the weak control strength of audio modality and the strong control strength of visual modality.
To address this issue, we introduce MegActor-Sigma: a mixed-modal conditional diffusion transformer (DiT), which can flexibly inject audio and visual modality control signals into portrait animation.
Specifically, we make substantial advancements over its predecessor, MegActor, by leveraging the promising model structure of DiT and integrating audio and visual conditions through advanced modules within the DiT framework.
To further achieve flexible combinations of mixed-modal control signals, we propose a Modality Decoupling Control" training strategy to balance the control strength between visual and audio modalities, along with the Amplitude Adjustment" inference strategy to freely regulate the motion amplitude of each modality.
Finally, to facilitate extensive studies in this field, we design several dataset evaluation metrics to filter out public datasets and solely use this filtered dataset for training.
Extensive experiments demonstrate the superiority of our approach in generating vivid portrait animations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PersonaLive! Expressive Portrait Image Animation for Live StreamingZhiyuan Li, Chi-Man Pun, Chen Fang, Jue Wang 等CVPR 2026 · 被引用 6 次
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute TransferHyunsoo Cha, Byungjun Kim, Hanbyul JooICLR 2026 · 被引用 2 次
- X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake DetectionYoungseo Kim, Kwan Yun, Seokhyeon Hong, Sihun Cha 等CVPR 2026 · 被引用 2 次
- Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking HeadsHaoyu Wang, Xiaozhe Xin, Xiaoyu Qin, Meiguang Jin 等AAAI 2026 · 被引用 1 次
- FG-Portrait: 3D Flow Guided Editable Portrait AnimationYating Xu, Yunqi Miao, Evangelos Ververas, Jiankang Deng 等CVPR 2026
它引用的顶会 Paper27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
相关 Paper
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsYuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu 等CVPR 2026 · 被引用 3 次
- DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid GuidanceYuxuan Luo, Zhengkun Rong, Lizhen Wang, Longhao Zhang 等ICCV 2025 · 被引用 5 次
- ExpPortrait: Expressive Portrait Generation via Personalized RepresentationJunyi Wang, Yudong Guo, Boyang Guo, Shengming Yang 等CVPR 2026
- X-Portrait: Expressive Portrait Animation with Hierarchical Motion AttentionYou Xie, Hongyi Xu, Guoxian Song, Chao Wang 等SIGGRAPH 2024 · 被引用 40 次
- RealPortrait: Realistic Portrait Animation with Diffusion TransformersZejun Yang, Huawei Wei, Zhisheng WangAAAI 2025 · 被引用 2 次
