Do LLMs Understand Dialogues? A Case Study on Dialogue Acts
Ayesha Qamar, Jonathan Tong, Ruihong Huang
摘要
Recent advancements in NLP, largely driven by Large Language Models (LLMs), have significantly improved performance on an array of tasks. However, Dialogue Act (DA) classification remains challenging, particularly in the fine-grained 50-class, multiparty setting. This paper investigates the root causes of LLMs’ poor performance in DA classification through a linguistically motivated analysis. We identify three key pre-tasks essential for accurate DA prediction: Turn Management , Communica-tive Function Identification , and Dialogue Structure Prediction . Our experiments reveal that LLMs struggle with these fundamental tasks, often failing to outperform simple rule-based baselines. Additionally, we establish a strong empirical correlation between errors in these pre-tasks and DA classification failures. A human study further highlights the significant gap between LLM and human-level dialogue understanding. These findings indicate that LLMs’ shortcomings in dialogue comprehension hinder their ability to accurately predict DAs, highlighting the need for improved dialogue-aware training approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon 等ACL 2024 · 被引用 14 次
- Zero-Shot Cross-Domain Dialogue State Tracking via Dual Low-Rank AdaptationXiang Luo, Zhiwen Tang, Jin Wang, Xuejie ZhangACL 2024 · 被引用 5 次
- ESCoT: Towards Interpretable Emotional Support Dialogue SystemsTenggan Zhang, Xinjie Zhang, Jinming Zhao, Li Zhou 等ACL 2024
- Towards Emotional Support Dialog SystemsSiyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour 等ACL 2021
相关 Paper
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He 等ACL 2024 · 被引用 35 次
- Evaluating the Effectiveness of Large Language Models in Establishing Conversational GroundingBiswesh Mohapatra, Manav Nitin Kapadnis, Laurent Romary, Justine CassellEMNLP 2024 · 被引用 1 次
- Probing LLMs for Multilingual Discourse Generalization Through a Unified Label SetFlorian Eichin, Yang Janet Liu, Barbara Plank, Michael A. HedderichACL 2025
- Masking Orchestration: Multi-Task Pretraining for Multi-Role Dialogue Representation LearningTianyi Wang, Yating Zhang, Xiaozhong Liu, Changlong Sun 等AAAI 2020 · 被引用 8 次
- Probing Task-Oriented Dialogue Representation from Language ModelsChien-Sheng Wu, Caiming XiongEMNLP 2020 · 被引用 20 次
