Lune

ACL2026Top-tier venue

ArgGenBench: Benchmarking the Complex Controlled Argument Generation Capability of Large Language Models

Bojun Jin, Jianzhu Bao, Yang Sun, Yice Zhang, Ruifeng Xu

2026Year

Abstract

Argument generation is a fundamental NLP task that aims to automatically produce persuasive arguments. Effective human argumentation is inherently complex and multifaceted, integrating argumentative strategies, appropriate styles, and adaptation to target audiences, etc. However, existing studies focus on limited control signals such as topic, stance, or key aspects, failing to capture this complexity. As LLMs advance, the lack of benchmarks evaluating multifaceted argumentative control becomes a critical bottleneck. To address this, we introduce ArgGenBench, a novel benchmark containing complex instructions that integrate multi-dimensional control, including topic, stance, length, style, strategy, audience, and key points. Extensive evaluation across 15 LLMs reveals significant limitations: even the best-performing model achieves only 42.7% win rate against human-verified references. These results highlight the challenge of controlled argument generation and establish ArgGenBench as a rigorous testbed for developing more capable systems.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5c68e458-3589-4a84-8201-72fb89d736ed

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines