MAViD: A Multimodal Framework for Audio-Visual Dialogue Understanding and Generation
Youxin Pang, Jiajun Liu, Lingfeng Tan, Yong Zhang, Feng Gao, Xiang Deng, Zhuoliang Kang, Xiaoming Wei, Yebin Liu
Abstract
We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The primary challenge of this task lies in effectively integrating understanding and generation capabilities, as well as achieving seamless multimodal audio-video fusion. To solve these problems, we propose a Conductor-Creator architecture that divides the dialogue system into two primary components. The Conductor is tasked with understanding, reasoning, and generating instructions by breaking them down into motion and speech components, thereby enabling fine-grained control over interactions. The Creator then delivers interactive responses based on these instructions. Furthermore, to address the difficulty of generating long videos with consistent identity, timbre, and tone using dual DiT structures, the Creator adopts a structure that combines autoregressive (AR) and diffusion models. The AR model is responsible for audio generation, while the diffusion model ensures high-quality video generation. Additionally, we propose a novel fusion module to enhance connections between contextually consecutive clips and modalities, enabling synchronized long-duration audio-visual content generation. Extensive experiments demonstrate that our framework can generate vivid and contextually coherent long-duration dialogue interactions and accurately interpret users' multimodal queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- MovieDreamer: Hierarchical Generation for Coherent Long Visual SequencesCanyu Zhao, Mingyu Liu, Wen Wang, Weihua Chen et al.ICLR 2025
- Ditailistener: Controllable High Fidelity Listener Video Generation with DiffusionMaksim Siniukov, Di Chang, Minh Tran, Hongkun Gong et al.ICCV 2025 · 1 citation
- UniVideo: Unified Understanding, Generation, and Editing for VideosCong Wei, Quande Liu, Zixuan Ye, Qiulin Wang et al.ICLR 2026 · 90 citations
- BindWeave: Subject-Consistent Video Generation via Cross-Modal IntegrationZhaoyang Li, Dongjun Qian, Kai Su, qishuai diao et al.ICLR 2026 · 23 citations
- JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and GenerationKai Liu, Jungang Li, Yuchong Sun, Shengqiong Wu et al.NeurIPS 2025 · 18 citations
