OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking
Zhongjian Wang, Peng Zhang, Jinwei Qi, Yuan Wang, Sheng Xu, Bang Zhang
摘要
Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video content from input text while emulating the speaking and facial movement styles of the target identity, including speech characteristics, head motion, and facial dynamics. Our framework adopts a dual-branch diffusion transformer (DiT) architecture, with one branch dedicated to audio generation and the other to video synthesis. At the shallow layers, cross-modal fusion modules are introduced to integrate information between the two modalities. In deeper layers, each modality is processed independently, with the generated audio decoded by a vocoder and the video rendered using a GAN-based high-quality visual renderer. Leveraging the in-context learning capability of DiT through a masked-infilling strategy, our model can simultaneously capture both audio and visual styles without requiring explicit style extraction modules. Thanks to the efficiency of the DiT backbone and the optimized visual renderer, OmniTalker achieves real-time inference at 25 FPS. To the best of our knowledge, OmniTalker is the first one-shot framework capable of jointly modeling speech and facial styles in real time. Extensive experiments demonstrate its superiority over existing methods in terms of generation quality, particularly in preserving style consistency and ensuring precise audio-video synchronization, all while maintaining efficient inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking HeadChang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou 等CVPR 2026 · 被引用 1 次
- Emotion-Conditioned Motion Sub-spaces with Flow Matching for Real-Time Audio-Driven Talking HeadsHaoyu Wang, Xiaozhe Xin, Xiaoyu Qin, Meiguang Jin 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeSicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang 等NeurIPS 2024 · 被引用 253 次
相关 Paper
- MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head GenerationSeyeon Kim, Siyoon Jin, Jihye Park, Kihong Kim 等AAAI 2025 · 被引用 12 次
- Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head SynthesisTianqi Li, Ruobing Zheng, Minghui Yang, Jingdong Chen 等ACM MM 2025 · 被引用 4 次
- Animate and Sound an ImageXihua Wang, Ruihua Song, Chongxuan Li, Xin Cheng 等CVPR 2025
- AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video GenerationKai Wang, Shijian Deng, Jing Shi, Dimitrios Hatzinakos 等ACM MM 2025 · 被引用 2 次
- READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head GenerationHaotian Wang, Yuzhe Weng, Jun Du, Haoran Xu 等AAAI 2026 · 被引用 1 次
