SyncDreamer: Controllable and Expressive Avatar Generation Beyond the Talking Head
Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia, Josef Kittler, Muhammad Awais
Abstract
Generating realistic and expressive audio-driven talking avatars remains a challenge in digital human synthesis. Existing methods often depend on intermediate representations for natural body motion, restricting flexibility and leading to visual distortions. Moreover, many approaches rely on discrete emotion labels to regulate expression. Such categorical supervision fails to capture the continuous and fine-grained speech dynamics (rhythm, energy, intensity), resulting in limited synchronization and emotionally shallow motion. To overcome these limitations, we present SyncDreamer, a unified diffusion Transformer framework that generates identity-preserving and emotionally expressive talking avatars from only a single image, speech audio, and text prompt. We propose a visual adapter with Attention Localization Loss to maintain identity fidelity, further incorporating an audio dynamics encoder for rhythm-and emotion-aware motion, and an RL-based Cross-Modal Prompt Enhancer grounding textual cues in visual context for fine-grained motion control. Extensive experiments on portrait and full-body benchmarks demonstrate state-of-the-art performance in realism, synchronization accuracy, and semantic controllability, establishing a scalable foundation for expressive digital avatars in interactive and creative applications. https: //fnazarieh.github.io/SyncDreamerWeb/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42c8ebce-29f6-47b2-9b5b-d37169c84805Builds on28
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion SynthesisMengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan et al.ACM MM 2025 · 9 citations
- AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsAggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer et al.ICCV 2025 · 3 citations
- Versatile Multimodal Controls for Expressive Talking Human AnimationZheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li et al.ACM MM 2025 · 2 citations
- AudioAvatar: Personalized Audio-driven Whole-body Talking AvatarsSeungeun Lee, SeungJun Moon, Hah Min Lew, Ji-Su Kang et al.CVPR 2026
- Expressive Talking Human from Single-Image with Imperfect PriorsJun Xiang, Yudong Guo, Leipeng Hu, Boyang Guo et al.ICCV 2025 · 3 citations
