How to Compare Things Properly? A Study of Argument Relevance in Comparative Question Answering
Irina Nikishina, Saba Anwar, Nikolay Dolgov, Maria Manina, Daria Ignatenko, Artem Shelmanov, Chris Biemann
摘要
Comparative Question Answering (CQA) lies at the intersection of Question Answering, Ar-gument Mining, and Summarization. It poses unique challenges due to the inherently subjective nature of many questions and the need to integrate diverse perspectives. Although the CQA task can be addressed using recently emerged instruction-following Large Language Models (LLMs), challenges such as hallucinations in their outputs and the lack of transparent argument provenance remain significant limitations. To address these challenges, we construct a manually curated dataset comprising arguments annotated with their relevance. These arguments are further used to answer comparative questions, enabling precise traceability and faithfulness. Furthermore, we define explicit criteria for an “ideal” comparison and introduce a benchmark for evaluating the outputs of various Retrieval-Augmented Generation (RAG) models with respect to argument relevance. All code and data are publicly released to support further research 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
- Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought MethodYiming Wang, Zhuosheng Zhang, Rui WangACL 2023 · 被引用 39 次
- Learning Opinion Summarizers by Selecting Informative ReviewsArthur Brazinskas, Mirella Lapata, Ivan TitovEMNLP 2021 · 被引用 26 次
- A Reality Check on Context Utilisation for Retrieval-Augmented GenerationLovisa Hagström, Sara Vera Marjanovic, Haeun Yu, Arnav Arora 等ACL 2025
- Towards Argument Mining for Social Good: A SurveyEva Maria Vecchi, Neele Falk, Iman Jundi, Gabriella LapesaACL 2021
相关 Paper
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language ModelsCheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu 等ACL 2024
- RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question AnsweringRujun Han, Yuhao Zhang, Peng Qi, Yumo Xu 等EMNLP 2024 · 被引用 10 次
- Structure Guided Retrieval-Augmented Generation for Factual QueriesMiao Xie, Xiao Zhang, Yi Li, Chunli LvACL 2026 · 被引用 1 次
- MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question AnsweringTeng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang 等EMNLP 2025 · 被引用 2 次
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
