CEDAR: A Chinese Evaluation Dataset for Computational Argumentation
Tian Lan, Jiang Li, Rong Yan, Feilong Bao, Weihua Wang, Guanglai Gao, Xiangdong Su
Abstract
Computational argumentation has received increasing attention in recent years. However, existing debate datasets neglect some important labels for argument mining, generation, and evaluation. Meanwhile, the lack of comprehensively annotated Chinese oral debate datasets hinders progress in this field. To address these gaps, we introduce a comprehensive Chinese Evaluation Dataset for Computational Argumentation, named CEDAR. Compared to previous datasets, CEDAR includes the essential labels of computational argumentation (claim, stance, evidence) and five additional crucial labels: rhetorical figures, debater roles, modal words, utterance time, and debate results. Moreover, it offers complete transcripts of each debate, including speeches from the Pro and Con sides. Thus, the proposed CEDAR not only supports common argument mining and generation tasks, but also provides resources for rhetorical figure detection, argument quality evaluation, and debate result prediction. This dataset covers 600 debates about 318 topics from Chinese debate competitions. Besides providing a dataset for research, we conduct experiments on common computational argument tasks and a novel task (rhetorical figure detection), in which we also evaluate LLMs. The experimental results highlight the challenging nature of the dataset. Our corpus is available at https://github.com/VelikayaScarlet/ CEDAR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 892cd3f5-c87d-4a20-9b89-6ee515c4bf87Builds on10
- Detecting Attackable Sentences in ArgumentsYohan Jo, Seojin Bang, Emaad A. Manzoor, Eduard H. Hovy et al.EMNLP 2020 · 25 citations
- Exploring the Potential of Large Language Models in Computational ArgumentationGuizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong BingACL 2024 · 8 citations
- ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue SummarizationXiutian Zhao, Ke Wang, Wei PengEMNLP 2023 · 5 citations
- Argue with Me Tersely: Towards Sentence-Level Counter-Argument GenerationJiayu Lin, Rong Ye, Meng Han, Qi Zhang et al.EMNLP 2023 · 2 citations
- The Moral Debater: A Study on the Computational Generation of Morally Framed ArgumentsMilad Alshomary, Roxanne El Baff, Timon Gurcke, Henning WachsmuthACL 2022
Related papers
- IAM: A Comprehensive and Large-Scale Dataset for Integrated Argument Mining TasksLiying Cheng, Lidong Bing, Ruidan He, Qian Yu et al.ACL 2022
- Hitting your MARQ: Multimodal ARgument Quality Assessment in Long Debate VideoMd. Kamrul Hasan, James Spann, Masum Hasan, Md. Saiful Islam et al.EMNLP 2021
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope et al.EMNLP 2025 · 1 citation
- Extracting Implicitly Asserted Propositions in ArgumentationYohan Jo, Jacky Visser, Chris Reed, Eduard H. HovyEMNLP 2020 · 2 citations
- SAD: A Large-Scale Strategic Argumentative Dialogue DatasetYongkang Liu, Jiayang Yu, Mingyang Wang, Yiqun Zhang et al.ACL 2026
