Beyond Recognising Entailment: Formalising Natural Language Inference from an Argumentative Perspective
Ameer Saadat-Yazdi, Nadin Kökciyan
Abstract
In argumentation theory, argument schemes are a characterisation of stereotypical patterns of inference. There has been little work done to develop computational approaches to identify these schemes in natural language. Moreover, advancements in recognizing textual entailment lack a standardized definition of inference, which makes it challenging to compare methods trained on different datasets and rely on the generalisability of their results. In this work, we propose a rigorous approach to align entailment recognition with argumentation theory. Wagemans' Periodic Table of Arguments (PTA), a taxonomy of argument schemes, provides the appropriate framework to unify these two fields. To operationalise the theoretical model, we introduce a tool to assist humans in annotating arguments according to the PTA. Beyond providing insights into non-expert annotator training, we present Kialo-PTA24, the first multi-topic dataset for the PTA. Finally, we benchmark the performance of pre-trained language models on various aspects of argument analysis. Our experiments show that the task of argument canonicalisation poses a significant challenge for state-of-the-art models, suggesting an inability to represent argumentative reasoning and a direction for future investigation.
-
We conduct an annotation study to rephrase natural language arguments into structured templates and provide insights into how to train non-expert annotators to perform this analysis.
-
We introduce ArgNotator, a tool that assists humans in annotating arguments according to the PTA.
-
We construct Kialo-PTA24 -the first multi-topic dataset of argument types annotated according to the PTA.
-
We compare the performance of state-of-theart models for two annotation subtasks. For the substance classification task, we benchmark the performance of a number of BERT-based models. For the argument canonicalisation task, we evaluate the performance of two large language models (FLAN-T5, LLAMA2) in both pre-trained and few-shot settings.
The dataset, experimental setup, annotation tool and training materials can all be found on GitLab. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ee46c23-c338-45fb-9b7c-159124e79534Builds on2
Related papers
- Mining Complex Patterns of Argumentative Reasoning in Natural Language DialogueRamon Ruiz-Dolz, Zlata Kikteva, John LawrenceACL 2025 · 2 citations
- Flee the Flaw: Annotating the Underlying Logic of Fallacious Arguments Through Templates and Slot-fillingIrfan Robbani, Paul Reisert, Surawat Pothong, Naoya Inoue et al.EMNLP 2024
- Leveraging pre-trained language models for linguistic analysis: A case of argument structure constructionsHakyung Sung, Kristopher KyleEMNLP 2024 · 2 citations
- How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field StudyAndreas Waldis, Yufang Hou, Iryna GurevychACL 2024
- ArgU: A Controllable Factual Argument GeneratorSougata Saha, Rohini K. SrihariACL 2023 · 4 citations
