Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets
Benjamin Schiller, Johannes Daxenberger, Andreas Waldis, Iryna Gurevych
Abstract
Topic-Dependent Argument Mining (TDAM), that is extracting and classifying argument components for a specific topic from large document sources, is an inherently difficult task for machine learning models and humans alike, as large TDAM datasets are rare and recognition of argument components requires expert knowledge. The task becomes even more difficult if it also involves stance detection of retrieved arguments. In this work, we investigate the effect of TDAM dataset composition in few-and zeroshot settings. Our findings show that, while fine-tuning is mandatory to achieve acceptable model performance, using carefully composed training samples and reducing the training sample size by up to almost 90% can still yield 95% of the maximum performance. This gain is consistent across three TDAM tasks on three different datasets. We also publish a new dataset 1 and code 2 for future benchmarking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0fbae8c-71db-4300-bc3d-f489d6428635Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
Related papers
- -Stance: A Large-Scale Real World Dataset of Stances in Legal ArgumentationAnkita Gupta, Douglas Rice, Brendan T. O'ConnorACL 2025
- EZ-STANCE: A Large Dataset for English Zero-Shot Stance DetectionChenye Zhao, Cornelia CarageaACL 2024
- Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic RepresentationsEmily Allaway, Kathleen R. McKeownEMNLP 2020 · 7 citations
- Unsupervised stance detection for arguments from consequencesJonathan Kobbe, Ioana Hulpus, Heiner StuckenschmidtEMNLP 2020 · 27 citations
- Fine-Grained Argument Unit Recognition and ClassificationDietrich Trautmann, Johannes Daxenberger, Christian Stab, Hinrich Schütze et al.AAAI 2020 · 70 citations
