ArgCMV: An Argument Summarization Benchmark for the LLM-era
Omkar Gurjar, Agam Goyal, Eshwar Chandrasekharan
Abstract
Key point extraction is an important task in argument summarization, which involves extracting high-level short summaries from arguments. Existing approaches for KP extraction have been mostly evaluated on the popular ArgKP21 dataset. In this paper, we highlight some of the major limitations of the ArgKP21 dataset and demonstrate the need for new benchmarks that are more representative of actual human conversations. Using SoTA large language models (LLMs), we curate a new argument key point extraction dataset called ArgCMV comprising of ∼ 12K arguments from actual online human debates spread across ∼ 3K topics. Our dataset exhibits higher complexity such as longer, coreferencing arguments, higher presence of subjective discourse units, and a larger range of topics over ArgKP21. We show that existing methods do not adapt well to ArgCMV and provide extensive benchmark results by experimenting with existing baselines and latest open source models. This work introduces a novel KP extraction dataset for long-context online discussions, setting the stage for the next generation of LLM-driven summarization research. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext babeb3a9-659d-4676-aaaa-c49b3754d8f8Cited by top-tier papers2
- PerSpectra: A Scalable and Configurable Pluralist Benchmark of Perspectives from ArgumentsShangrui Nie, Kian Omoomi, Lucie Flek, Zhixue Zhao et al.ICLR 2026 · 3 citations
- Needling Through the Threads: A Visualization Tool for Navigating Threaded Online DiscussionsYijun Liu, Frederick Choi, Eshwar ChandrasekharanCHI 2026 · 1 citation
Builds on14
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 173 citations
- A Large-Scale Dataset for Argument Quality Ranking: Construction and AnalysisShai Gretz, Roni Friedman, Edo Cohen-Karlik, Assaf Toledo et al.AAAI 2020 · 148 citations
- SolutionChat: Real-time Moderator Support for Chat-based Structured DiscussionSung-Chul Lee, Jaeyoon Song, Eun-Young Ko, Seongho Park et al.CHI 2020 · 47 citations
- Proactive Moderation of Online Discussions: Existing Practices and the Potential for Algorithmic SupportCharlotte Schluger, Jonathan P. Chang, Cristian Danescu-Niculescu-Mizil, Karen LevyCSCW 2022 · 41 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
Related papers
- From Arguments to Key Points: Towards Automatic Argument SummarizationRoy Bar-Haim, Lilach Eden, Roni Friedman, Yoav Kantor et al.ACL 2020 · 5 citations
- ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument MiningAlexander R. Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang et al.ACL 2021
- Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning ExtractionWenxuan Liu, Zixuan Li, Long Bai, Yuxin Zuo et al.EMNLP 2025
- How to Compare Things Properly? A Study of Argument Relevance in Comparative Question AnsweringIrina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina et al.ACL 2025
- Exploring the Potential of Large Language Models in Computational ArgumentationGuizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong BingACL 2024 · 8 citations
