Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations
Nicolò Penzo, Maryam Sajedinia, Bruno Lepri, Sara Tonelli, Marco Guerini
摘要
Assessing the performance of systems to classify Multi-Party Conversations (MPC) is challenging due to the interconnection between linguistic and structural characteristics of conversations. Conventional evaluation methods often overlook variances in model behavior across different levels of structural complexity on interaction graphs. In this work, we propose a methodological pipeline to investigate model performance across specific structural attributes of conversations. As a proof of concept we focus on Response Selection and Addressee Recognition tasks, to diagnose model weaknesses. To this end, we extract representative diagnostic subdatasets with a fixed number of users and a good structural variety from a large and open corpus of online MPCs. We further frame our work in terms of data minimization, avoiding the use of original usernames to preserve privacy, and propose alternatives to using original text messages. Results show that response selection relies more on the textual content of conversations, while addressee recognition requires capturing their structural dimension. Using an LLM in a zero-shot setting, we further highlight how sensitivity to prompt variations is task-dependent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent SafetySanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang 等ICLR 2026 · 被引用 75 次
- Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with ConstraintsNicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas 等AAAI 2026 · 被引用 3 次
它引用的顶会 Paper7
- ProPILE: Probing Privacy Leakage in Large Language ModelsSiwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri 等NeurIPS 2023 · 被引用 229 次
- Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsJiuding Sun, Chantal Shaib, Byron C. WallaceICLR 2024 · 被引用 75 次
- Response Selection for Multi-Party Conversations with Dynamic Topic TrackingWeishi Wang, Steven C. H. Hoi, Shafiq R. JotyEMNLP 2020 · 被引用 41 次
- Multi-turn Response Selection using Dialogue Dependency RelationsQi Jia, Yizhu Liu, Siyu Ren, Kenny Q. Zhu 等EMNLP 2020 · 被引用 31 次
- GIFT: Graph-Induced Fine-Tuning for Multi-Party Conversation UnderstandingJia-Chen Gu, Zhenhua Ling, Quan Liu, Cong Liu 等ACL 2023 · 被引用 3 次
相关 Paper
- MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation UnderstandingJia-Chen Gu, Chongyang Tao, Zhen-Hua Ling, Can Xu 等ACL 2021
- MADNet: Maximizing Addressee Deduction Expectation for Multi-Party Conversation GenerationJia-Chen Gu, Chao-Hong Tan, Caiyuan Chu, Zhen-Hua Ling 等EMNLP 2023
- EM Pre-training for Multi-party Dialogue Response GenerationYiyang Li, Hai ZhaoACL 2023 · 被引用 9 次
- HeterMPC: A Heterogeneous Graph Neural Network for Response Generation in Multi-Party ConversationsJia-Chen Gu, Chao-Hong Tan, Chongyang Tao, Zhen-Hua Ling 等ACL 2022
- PII-Bench: Evaluating Query-Aware Privacy Protection SystemsHao Shen, Zhouhong Gu, Haokai Hong, Weili Han 等ACL 2026
