MegActor-Sigma: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
Shurong Yang, Huadong Li, Juhao Wu, Minhao Jing, Linze Li, Renhe Ji, Jiajun Liang, Haoqiang Fan, Jin Wang
Abstract
Diffusion models have demonstrated superior performance in portrait animation.
However, current approaches relied on either visual or audio modality to control character movements, failing to exploit the potential of mixed-modal control.
This challenge arises from the difficulty in balancing the weak control strength of audio modality and the strong control strength of visual modality.
To address this issue, we introduce MegActor-Sigma: a mixed-modal conditional diffusion transformer (DiT), which can flexibly inject audio and visual modality control signals into portrait animation.
Specifically, we make substantial advancements over its predecessor, MegActor, by leveraging the promising model structure of DiT and integrating audio and visual conditions through advanced modules within the DiT framework.
To further achieve flexible combinations of mixed-modal control signals, we propose a Modality Decoupling Control" training strategy to balance the control strength between visual and audio modalities, along with the Amplitude Adjustment" inference strategy to freely regulate the motion amplitude of each modality.
Finally, to facilitate extensive studies in this field, we design several dataset evaluation metrics to filter out public datasets and solely use this filtered dataset for training.
Extensive experiments demonstrate the superiority of our approach in generating vivid portrait animations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd5f5bb3-c448-4957-82e6-7ef34b742d7fCited by top-tier papers5
- PersonaLive! Expressive Portrait Image Animation for Live StreamingZhiyuan Li, Chi-Man Pun, Chen Fang, Jue Wang et al.CVPR 2026 · 6 citations
- Durian: Dual Reference Image-Guided Portrait Animation with Attribute TransferHyunsoo Cha, Byungjun Kim, Hanbyul JooICLR 2026 · 2 citations
- X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake DetectionYoungseo Kim, Kwan Yun, Seokhyeon Hong, Sihun Cha et al.CVPR 2026 · 2 citations
- Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking HeadsHaoyu Wang, Xiaozhe Xin, Xiaoyu Qin, Meiguang Jin et al.AAAI 2026 · 1 citation
- FG-Portrait: 3D Flow Guided Editable Portrait AnimationYating Xu, Yunqi Miao, Evangelos Ververas, Jiankang Deng et al.CVPR 2026
Builds on27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel et al.ICCV 2023 · 800 citations
Related papers
- DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsYuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu et al.CVPR 2026 · 3 citations
- DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid GuidanceYuxuan Luo, Zhengkun Rong, Lizhen Wang, Longhao Zhang et al.ICCV 2025 · 5 citations
- ExpPortrait: Expressive Portrait Generation via Personalized RepresentationJunyi Wang, Yudong Guo, Boyang Guo, Shengming Yang et al.CVPR 2026
- X-Portrait: Expressive Portrait Animation with Hierarchical Motion AttentionYou Xie, Hongyi Xu, Guoxian Song, Chao Wang et al.SIGGRAPH 2024 · 40 citations
- RealPortrait: Realistic Portrait Animation with Diffusion TransformersZejun Yang, Huawei Wei, Zhisheng WangAAAI 2025 · 2 citations
