LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming
Jingsheng Gao, Yixin Lian, Ziyi Zhou, Yuzhuo Fu, Baoyuan Wang
Abstract
Open-domain dialogue systems have made promising progress in recent years. While the state-of-the-art dialogue agents are built upon large-scale text-based social media data and large pre-trained models, there is no guarantee these agents could also perform well in fast-growing scenarios, such as live streaming, due to the bounded transferability of pretrained models and biased distributions of public datasets from Reddit and Weibo, etc. To improve the essential capability of responding and establish a benchmark in the live opendomain scenario, we introduce the LiveChat dataset, composed of 1.33 million real-life Chinese dialogues with almost 3800 average sessions across 351 personas and fine-grained profiles for each persona. LiveChat is automatically constructed by processing numerous live videos on the Internet and naturally falls within the scope of multi-party conversations, where the issues of Who says What to Whom should be considered. Therefore, we target two critical tasks of response modeling and addressee recognition and propose retrieval-based baselines grounded on advanced techniques. Experimental results have validated the positive effects of leveraging persona profiles and larger average sessions per persona. In addition, we also benchmark the transferability of advanced generation-based models on LiveChat and pose some future directions for current challenges. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92242f90-1efe-4e84-99a1-424c2712710eCited by top-tier papers8
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- LiveStar: Live Streaming Assistant for Real-World Online Video UnderstandingZhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang et al.NeurIPS 2025 · 26 citations
- If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMsSiqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang et al.ACL 2026 · 4 citations
- Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional WorksXinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin et al.EMNLP 2024 · 2 citations
- InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological InterviewsXintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan et al.ACL 2024
Builds on5
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 329 citations
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with LanguageAndy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski et al.ICLR 2023 · 171 citations
- Towards Persona-Based Empathetic Conversational ModelsPeixiang Zhong, Chen Zhang, Hao Wang, Yong Liu et al.EMNLP 2020 · 112 citations
- MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation UnderstandingJia-Chen Gu, Chongyang Tao, Zhen-Hua Ling, Can Xu et al.ACL 2021
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding et al.ACL 2022
Related papers
- MedDialog: Large-scale Medical Dialogue DatasetsGuangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang et al.EMNLP 2020 · 163 citations
- Re³Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingJiaxin Wen, Hao Zhou, Jian Guan, Jie Zhou et al.EMNLP 2023 · 2 citations
- Dual Task Framework for Improving Persona-Grounded Dialogue DatasetMinju Kim, Beong-woo Kwak, Youngwook Kim, Hong-in Lee et al.AAAI 2022 · 8 citations
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 4 citations
- CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog EvaluationYinpei Dai, Wanwei He, Bowen Li, Yuchuan Wu et al.EMNLP 2022 · 6 citations
