Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim, Joanna Hong, Jeong Hun Yeo, Yong Man Ro
摘要
In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system without relying on intermediate text. To this end, we newly introduce MultiDialog, the first large-scale multimodal (i.e., audio and visual) spoken dialogue corpus containing 340 hours of approximately 9,000 dialogues, recorded based on the open domain dialogue dataset, TopicalChat. The MultiDialog contains parallel audio-visual recordings of conversation partners acting according to the given script with emotion annotations, which we expect to open up research opportunities in multimodal synthesis. Our Face-to-Face spoken dialogue model incorporates a textually pretrained large language model and adapts it into the audio-visual spoken dialogue domain by incorporating speech-text joint pretraining. Through extensive experiments, we validate the effectiveness of our model in facilitating a face-to-face conversation. Demo and data are available at https://multidialog.github.io and https://huggingface.co/datasets/ IVLLab/MultiDialog , respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and GenerationKai Liu, Jungang Li, Yuchong Sun, Shengqiong Wu 等NeurIPS 2025 · 被引用 18 次
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic InteractionsCheng Luo, Jianghui Wang, Bing Li, Siyang Song 等NeurIPS 2025 · 被引用 4 次
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis 等ICCV 2025 · 被引用 3 次
- UniTalker: Conversational Speech-Visual SynthesisYifan Hu, Rui Liu, Yi Ren, Xiang Yin 等ACM MM 2025 · 被引用 2 次
- When Words Smile: Generating Diverse Emotional Facial Expressions from TextHaidong Xu, Meishan Zhang, Hao Ju, Zhedong Zheng 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji 等ICML 2024 · 被引用 786 次
相关 Paper
- Making Visual Dialogue More Engaging: A New Task, Method, and MetricGuanghui Ye, Huan Zhao, Yingxue Gao, Zhixue Zhao 等AAAI 2026
- MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain ConversationJiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao 等ACL 2023 · 被引用 20 次
- Scaling Speech-Text Pre-training with Synthetic Interleaved DataAohan Zeng, Zhengxiao Du, Mingdao Liu, Lei Zhang 等ICLR 2025
- ContextFace: Generating Facial Expressions from Emotional ContextsMin-Jung Kim, Minsang Kim, Seung Jun BaekICCV 2025 · 被引用 1 次
- Paralinguistics-Aware Speech-Empowered Large Language Models for Natural ConversationHeeseung Kim, Soonshin Seo, Kyeongseok Jeong, Ohsung Kwon 等NeurIPS 2024 · 被引用 28 次
