PACHAT: Persona-Aware Speech Assistant for Multi-party Dialogue
Dongjie Fu, Xize Cheng, Linjun Li, Xiaoda Yang, Lujia Yang, Tao Jin
Abstract
Extensive research on LLM-based spoken dialogue systems has significantly advanced the development of intelligent voice assistants. However, the integration of role information within speech remains an underexplored area, limiting its application in real-world scenarios, particularly in multi-party dialogue settings. With the growing demand for personalization, voice assistants that can recognize and remember users establish a deeper connection with them. We focus on enabling LLMs with speaker-awareness capabilities and enhancing their understanding of character settings through synthetic data to generate contextually appropriate responses. We introduce Persona-Dialogue, the first large-scale multi-party spoken dialogue dataset that incorporates speaker profiles. Based on this dataset, we propose PAChat, an architecture that simultaneously models both linguistic content and speaker features, allowing LLMs to map character settings to speaker identities in speech. Through extensive experiments, we demonstrate that PAChat successfully achieves speaker-specific responses, character understanding, and the generation of targeted replies in multi-party dialogue scenarios, surpassing existing spoken dialogue systems. For more details, please visit our demo page at https: //persona-dialogue.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1412c1f-59bf-429e-9ada-f71c782eaf58Cited by top-tier papers5
- MARS-Sep: Multimodal-Aligned Reinforced Sound SeparationZihan Zhang, Xize Cheng, Zhennan Jiang, Dongjie Fu et al.ICLR 2026 · 2 citations
- Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across DomainsFangming Feng, Sihang Cai, Zequn Xie, Yangyang Wu et al.AAAI 2026 · 1 citation
- Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-SpeechFangming Feng, Dongjie Fu, Zequn Xie, Yu Zhang et al.ACL 2026
- SpatialLogic-Bench: A Diagnostic Benchmark for Task-Oriented Spatiotemporal ReasoningXiaoda Yang, Shenzhou Gao, Can Wang, Jiahe Zhang et al.AAAI 2026
- PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio PerceptionTianwei Lan, Jiaqi Wu, Zeming Liu, Zhaoxin Fan et al.ACL 2026
Builds on13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and RecognitionXize Cheng, Tao Jin, Rongjie Huang, Linjun Li et al.ICCV 2023 · 30 citations
- Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken ConversationsGuan-Ting Lin, Cheng-Han Chiang, Hung-yi LeeACL 2024 · 15 citations
Related papers
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 4 citations
- Speaker Verification in Agent-generated ConversationsYizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang et al.ACL 2024
- A Pre-Training Based Personalized Dialogue Generation Model with Persona-Sparse DataYinhe Zheng, Rongsheng Zhang, Minlie Huang, Xiaoxi MaoAAAI 2020 · 173 citations
- V-VAE: A Variational Auto Encoding Framework Towards Fine-Grained Control over Human-Like ChatQi Lin, Weikai Xu, Lisi Chen, Bin DaiEMNLP 2025
- Learning to Memorize Entailment and Discourse Relations for Persona-Consistent DialoguesRuijun Chen, Jin Wang, Liang-Chih Yu, Xuejie ZhangAAAI 2023 · 32 citations
