How to Compare Things Properly? A Study of Argument Relevance in Comparative Question Answering
Irina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina, Daria Ignatenko, Artem Shelmanov, Chris Biemann
Abstract
Comparative Question Answering (CQA) lies at the intersection of Question Answering, Ar-gument Mining, and Summarization. It poses unique challenges due to the inherently subjective nature of many questions and the need to integrate diverse perspectives. Although the CQA task can be addressed using recently emerged instruction-following Large Language Models (LLMs), challenges such as hallucinations in their outputs and the lack of transparent argument provenance remain significant limitations. To address these challenges, we construct a manually curated dataset comprising arguments annotated with their relevance. These arguments are further used to answer comparative questions, enabling precise traceability and faithfulness. Furthermore, we define explicit criteria for an “ideal” comparison and introduce a benchmark for evaluating the outputs of various Retrieval-Augmented Generation (RAG) models with respect to argument relevance. All code and data are publicly released to support further research 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e01d566-0d64-42c5-b026-25f31155547dBuilds on5
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
- Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought MethodYiming Wang, Zhuosheng Zhang, Rui WangACL 2023 · 39 citations
- Learning Opinion Summarizers by Selecting Informative ReviewsArthur Brazinskas, Mirella Lapata, Ivan TitovEMNLP 2021 · 26 citations
- A Reality Check on Context Utilisation for Retrieval-Augmented GenerationLovisa Hagström, Sara Vera Marjanovic, Haeun Yu, Arnav Arora et al.ACL 2025
- Towards Argument Mining for Social Good: A SurveyEva Maria Vecchi, Neele Falk, Iman Jundi, Gabriella LapesaACL 2021
Related papers
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language ModelsCheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu et al.ACL 2024
- RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question AnsweringRujun Han, Yuhao Zhang, Peng Qi, Yumo Xu et al.EMNLP 2024 · 10 citations
- Structure Guided Retrieval-Augmented Generation for Factual QueriesMiao Xie, Xiao Zhang, Yi Li, Chunli LvACL 2026 · 1 citation
- MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question AnsweringTeng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang et al.EMNLP 2025 · 2 citations
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
