ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language Models
Bojun Jin, Jianzhu Bao, Yang Sun, Yice Zhang, Ruifeng Xu
摘要
Argument generation is a fundamental NLP task that aims to automatically produce persuasive arguments. Effective human argumentation is inherently complex and multifaceted, integrating argumentative strategies, appropriate styles, and adaptation to target audiences, etc. However, existing studies focus on limited control signals such as topic, stance, or key aspects, failing to capture this complexity. As LLMs advance, the lack of benchmarks evaluating multifaceted argumentative control becomes a critical bottleneck. To address this, we introduce ArgGenBench, a novel benchmark containing complex instructions that integrate multi-dimensional control, including topic, stance, length, style, strategy, audience, and key points. Extensive evaluation across 15 LLMs reveals significant limitations: even the best-performing model achieves only 42.7% win rate against human-verified references. These results highlight the challenge of controlled argument generation and establish ArgGenBench as a rigorous testbed for developing more capable systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- DialoGraph: Incorporating Interpretable Strategy-Graph Networks into Negotiation DialoguesRishabh Joshi, Vidhisha Balachandran, Shikhar Vashishth, Alan W. Black 等ICLR 2021 · 被引用 39 次
- Improving Multi-turn Emotional Support Dialogue Generation with Lookahead Strategy PlanningYi Cheng, Wenge Liu, Wenjie Li, Jiashuo Wang 等EMNLP 2022 · 被引用 32 次
- Weakly-Supervised Hierarchical Models for Predicting Persuasive Strategies in Good-faith Textual RequestsJiaao Chen, Diyi YangAAAI 2021 · 被引用 25 次
- MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang 等EMNLP 2024 · 被引用 14 次
- A Synthetic Data Generation Framework for Grounded DialoguesJianzhu Bao, Rui Wang, Yasheng Wang, Aixin Sun 等ACL 2023 · 被引用 11 次
相关 Paper
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationNoy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope 等EMNLP 2025 · 被引用 1 次
- ArgU: A Controllable Factual Argument GeneratorSougata Saha, Rohini K. SrihariACL 2023 · 被引用 4 次
- Exploring the Potential of Large Language Models in Computational ArgumentationGuizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong BingACL 2024 · 被引用 8 次
- Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic DifferencesArkadiusz Modzelewski, Pawel Golik, Anna Kolos, Giovanni Da San MartinoACL 2026 · 被引用 1 次
- ReFF: Reinforcing Format Faithfulness in Language Models Across Varied TasksJiashu Yao, Heyan Huang, Zeming Liu, Haoyu Wen 等AAAI 2025 · 被引用 1 次
