Exploring the Potential of Large Language Models in Computational Argumentation
Guizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong Bing
摘要
Computational argumentation has become an essential tool in various domains, including law, public policy, and artificial intelligence. It is an emerging research field in natural language processing that attracts increasing attention. Research on computational argumentation mainly involves two types of tasks: argument mining and argument generation. As large language models (LLMs) have demonstrated impressive capabilities in understanding context and generating natural language, it is worthwhile to evaluate the performance of LLMs on diverse computational argumentation tasks. This work aims to embark on an assessment of LLMs, such as ChatGPT, Flan models, and LLaMA2 models, in both zero-shot and few-shot settings. We organize existing tasks into six main categories and standardize the format of fourteen openly available datasets. In addition, we present a new benchmark dataset on counter speech generation that aims to holistically evaluate the end-to-end performance of LLMs on argument mining and argument generation. Extensive experiments show that LLMs exhibit commendable performance across most of the datasets, demonstrating their capabilities in the field of argumentation. Our analysis offers valuable suggestions for evaluating computational argumentation and its integration with LLMs in future research endeavors. 1 * Equal contribution. Guizhen Chen is under the Joint PhD Program between Alibaba and Nanyang Technological University. † Liying Cheng is the corresponding author. 1 Our data and code implementation are released at https://github.com/DAMO-NLP-SG/LLM-argumentation .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Argumentative Large Language Models for Explainable and Contestable Claim VerificationGabriel Freedman, Adam Dejl, Deniz Gorur, Xiang Yin 等AAAI 2025 · 被引用 31 次
- Harnessing Toulmin's theory for zero-shot argument explicationAnkita Gupta, Ethan Zuckerman, Brendan T. O'ConnorACL 2024 · 被引用 4 次
- Exit Stories: Using Reddit Self-Disclosures to Understand Disengagement from Problematic CommunitiesShruti PhadkeCSCW 2025 · 被引用 2 次
- Language is Scary when Over-Analyzed: Unpacking Implied Misogynistic Reasoning with Argumentation Theory-Driven PromptsArianna Muti, Federico Ruggeri, Khalid Al-Khatib, Alberto Barrón-Cedeño 等EMNLP 2024 · 被引用 1 次
- Evaluation and Facilitation of Online Discussions in the LLM Era: A SurveyKaterina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé 等EMNLP 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- UL2: Unifying Language Learning ParadigmsYi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia 等ICLR 2023 · 被引用 97 次
- Improving Neural Cross-Lingual Abstractive Summarization via Employing Optimal Transport Distance for Knowledge DistillationThong Thanh Nguyen, Anh Tuan LuuAAAI 2022 · 被引用 46 次
相关 Paper
- Argue with Me Tersely: Towards Sentence-Level Counter-Argument GenerationJiayu Lin, Rong Ye, Meng Han, Qi Zhang 等EMNLP 2023 · 被引用 2 次
- ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language ModelsBojun Jin, Jianzhu Bao, Yang Sun, Yice Zhang 等ACL 2026
- CEDAR: A Chinese Evaluation Dataset for Computational ArgumentationTian Lan, Jiang Li, Rong Yan, Feilong Bao 等ACL 2026
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope 等EMNLP 2025 · 被引用 1 次
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang 等AAAI 2024 · 被引用 57 次
