Argue with Me Tersely: Towards Sentence-Level Counter-Argument Generation
Jiayu Lin, Rong Ye, Meng Han, Qi Zhang, Ruofei Lai, Xinyu Zhang, Zhao Cao, Xuanjing Huang, Zhongyu Wei
Abstract
Counter-argument generation-a captivating area in computational linguistics-seeks to craft statements that offer opposing views. While most research has ventured into paragraph-level generation, sentence-level counter-argument generation beckons with its unique constraints and brevity-focused challenges. Furthermore, the diverse nature of counter-arguments poses challenges for evaluating model performance solely based on ngram-based metrics. In this paper, we present the ArgTersely benchmark for sentence-level counter-argument generation, drawing from a manually annotated dataset from the Change-MyView debate forum 1 . We also propose Arg-LlaMA for generating high-quality counterargument. For better evaluation, we trained a BERT-based evaluator Arg-Judge with human preference data. We conducted comparative experiments involving various baselines such as LlaMA, Alpaca, GPT-3, and others. The results show the competitiveness of our proposed framework and evaluator in counter-argument generation tasks. Code and data are available at https://github.com/ amazingljy1206/ArgTersely .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc842ed8-c3c5-4ab5-970d-ef0d5232a731Cited by top-tier papers9
- Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech CounteringHelena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio et al.EMNLP 2024 · 2 citations
- AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive ContextsEsra Dönmez, Maximilian Maurer, Gabriella Lapesa, Agnieszka FalenskaEMNLP 2025 · 1 citation
- Plan Dynamically, Express Rhetorically: A Debate-Driven Rhetorical Framework for Argumentative WritingXueguan Zhao, Wenpeng Lu, Chaoqun Zheng, Weiyu Zhang et al.EMNLP 2025 · 1 citation
- A Multi-persona Framework for Argument Quality AssessmentBojun Jin, Jianzhu Bao, Yufang Hou, Yang Sun et al.ACL 2025
- ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language ModelsBojun Jin, Jianzhu Bao, Yang Sun, Yice Zhang et al.ACL 2026
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- Exploring the Potential of Large Language Models in Computational ArgumentationGuizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong BingACL 2024 · 8 citations
- CLOMO: Counterfactual Logical Modification with Large Language ModelsYinya Huang, Ruixin Hong, Hongming Zhang, Wei Shao et al.ACL 2024 · 2 citations
- AEG: Argumentative Essay Generation via A Dual-Decoder Model with Content PlanningJianzhu Bao, Yasheng Wang, Yitong Li, Fei Mi et al.EMNLP 2022 · 1 citation
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope et al.EMNLP 2025 · 1 citation
- Contextual Interaction for Argument Post Quality AssessmentYiran Wang, Xuanang Chen, Ben He, Le SunEMNLP 2023 · 4 citations
