ArgCMV: An Argument Summarization Benchmark for the LLM-era
Omkar Gurjar, Agam Goyal, Eshwar Chandrasekharan
摘要
Key point extraction is an important task in argument summarization, which involves extracting high-level short summaries from arguments. Existing approaches for KP extraction have been mostly evaluated on the popular ArgKP21 dataset. In this paper, we highlight some of the major limitations of the ArgKP21 dataset and demonstrate the need for new benchmarks that are more representative of actual human conversations. Using SoTA large language models (LLMs), we curate a new argument key point extraction dataset called ArgCMV comprising of ∼ 12K arguments from actual online human debates spread across ∼ 3K topics. Our dataset exhibits higher complexity such as longer, coreferencing arguments, higher presence of subjective discourse units, and a larger range of topics over ArgKP21. We show that existing methods do not adapt well to ArgCMV and provide extensive benchmark results by experimenting with existing baselines and latest open source models. This work introduces a novel KP extraction dataset for long-context online discussions, setting the stage for the next generation of LLM-driven summarization research. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PerSpectra: A Scalable and Configurable Pluralist Benchmark of Perspectives from ArgumentsShangrui Nie, Kian Omoomi, Lucie Flek, Zhixue Zhao 等ICLR 2026 · 被引用 3 次
- Needling Through the Threads: A Visualization Tool for Navigating Threaded Online DiscussionsYijun Liu, Frederick Choi, Eshwar ChandrasekharanCHI 2026 · 被引用 1 次
它引用的顶会 Paper14
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 被引用 173 次
- A Large-Scale Dataset for Argument Quality Ranking: Construction and AnalysisShai Gretz, Roni Friedman, Edo Cohen-Karlik, Assaf Toledo 等AAAI 2020 · 被引用 148 次
- SolutionChat: Real-time Moderator Support for Chat-based Structured DiscussionSung-Chul Lee, Jaeyoon Song, Eun-Young Ko, Seongho Park 等CHI 2020 · 被引用 47 次
- Proactive Moderation of Online Discussions: Existing Practices and the Potential for Algorithmic SupportCharlotte Schluger, Jonathan P. Chang, Cristian Danescu-Niculescu-Mizil, Karen LevyCSCW 2022 · 被引用 41 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
相关 Paper
- From Arguments to Key Points: Towards Automatic Argument SummarizationRoy Bar-Haim, Lilach Eden, Roni Friedman, Yoav Kantor 等ACL 2020 · 被引用 5 次
- ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument MiningAlexander R. Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang 等ACL 2021
- Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning ExtractionWenxuan Liu, Zixuan Li, Long Bai, Yuxin Zuo 等EMNLP 2025
- How to Compare Things Properly? A Study of Argument Relevance in Comparative Question AnsweringIrina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina 等ACL 2025
- Exploring the Potential of Large Language Models in Computational ArgumentationGuizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong BingACL 2024 · 被引用 8 次
