LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming
Jingsheng Gao, Yixin Lian, Ziyi Zhou, Yuzhuo Fu, Baoyuan Wang
摘要
Open-domain dialogue systems have made promising progress in recent years. While the state-of-the-art dialogue agents are built upon large-scale text-based social media data and large pre-trained models, there is no guarantee these agents could also perform well in fast-growing scenarios, such as live streaming, due to the bounded transferability of pretrained models and biased distributions of public datasets from Reddit and Weibo, etc. To improve the essential capability of responding and establish a benchmark in the live opendomain scenario, we introduce the LiveChat dataset, composed of 1.33 million real-life Chinese dialogues with almost 3800 average sessions across 351 personas and fine-grained profiles for each persona. LiveChat is automatically constructed by processing numerous live videos on the Internet and naturally falls within the scope of multi-party conversations, where the issues of Who says What to Whom should be considered. Therefore, we target two critical tasks of response modeling and addressee recognition and propose retrieval-based baselines grounded on advanced techniques. Experimental results have validated the positive effects of leveraging persona profiles and larger average sessions per persona. In addition, we also benchmark the transferability of advanced generation-based models on LiveChat and pose some future directions for current challenges. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- LiveStar: Live Streaming Assistant for Real-World Online Video UnderstandingZhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang 等NeurIPS 2025 · 被引用 26 次
- If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMsSiqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang 等ACL 2026 · 被引用 4 次
- Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional WorksXinfeng Yuan, Siyu Yuan, Yuhan Cui, Tianhe Lin 等EMNLP 2024 · 被引用 2 次
- InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological InterviewsXintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan 等ACL 2024
它引用的顶会 Paper5
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 被引用 329 次
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with LanguageAndy Zeng, Maria Attarian, Brian Ichter, Krzysztof Marcin Choromanski 等ICLR 2023 · 被引用 171 次
- Towards Persona-Based Empathetic Conversational ModelsPeixiang Zhong, Chen Zhang, Hao Wang, Yong Liu 等EMNLP 2020 · 被引用 112 次
- MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation UnderstandingJia-Chen Gu, Chongyang Tao, Zhen-Hua Ling, Can Xu 等ACL 2021
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding 等ACL 2022
相关 Paper
- MedDialog: Large-scale Medical Dialogue DatasetsGuangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang 等EMNLP 2020 · 被引用 163 次
- Re³Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingJiaxin Wen, Hao Zhou, Jian Guan, Jie Zhou 等EMNLP 2023 · 被引用 2 次
- Dual Task Framework for Improving Persona-Grounded Dialogue DatasetMinju Kim, Beong-woo Kwak, Youngwook Kim, Hong-in Lee 等AAAI 2022 · 被引用 8 次
- MPCHAT: Towards Multimodal Persona-Grounded ConversationJaewoo Ahn, Yeda Song, Sangdoo Yun, Gunhee KimACL 2023 · 被引用 4 次
- CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog EvaluationYinpei Dai, Wanwei He, Bowen Li, Yuchuan Wu 等EMNLP 2022 · 被引用 6 次
