Benchmarking LLM's Capability in Reasoning over Conflicting Web References
Yizhen Yuan, Rui Kong, Dongze Li, Yuanchun Li, Yunxin Liu
摘要
Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have become a dominant framework for building intelligent assistants. In real-world applications such as ChatGPT with web search, the retrieved document often comes from diverse, potentially unreliable sources and may contain inconsistent claims. Unlike traditional search engines that rely on users to manually compare information, LLM-based systems typically feed all retrieved content into the model's context, requiring LLMs to autonomously identify, differentiate, and reason over conflicting viewpoints. Unlike mainstream LLM evaluation tasks like math and code generation that are primarily focused on reasoning with factual context, question-answering with multi-source references requires fundamentally different capabilities to identify and reason over knowledge contradictions. In this paper, we introduce CONFRAG, a benchmark for evaluating LLMs' reasoning capability over real-world conflicting documents retrieved from the web. It consists of 1,814 real-world questions, each paired with an average of 9.58 retrieved paragraphs from heterogeneous online sources. A total of 57.2% of the questions exhibit explicit contradictions. We further propose three structured evaluation tasks, answer clustering, answer coverage, and reason coverage, to quantify a model's ability to organize and explain contradictory content. Experiments with state-of-the-art models such as GPT-4.1 and Claude-3-7-Sonnet reveal substantial performance gaps, highlighting the need for more targeted research in contradiction-aware question answering. To the best of our knowledge, CONFRAG is the first benchmark specifically designed to evaluate contradiction-aware reasoning on real-world long web documents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- Synchromesh: Reliable Code Generation from Pre-trained Language ModelsGabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari 等ICLR 2022 · 被引用 200 次
- GEO: Generative Engine OptimizationPranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan 等KDD 2024 · 被引用 23 次
相关 Paper
- Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and MethodTianzhe Zhao, Jiaoyan Chen, Shuxiu Zhang, Haiping Zhu 等SIGIR 2026
- Conflict-Aware RAG: Multi-Stage Learning with Conflict Signals for Robust Retrieval-Augmented GenerationHaiyan Wu, Chenchen Wang, Chaoqun Sun, Chengxiong Lu 等WWW 2026
- MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question AnsweringTeng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang 等EMNLP 2025 · 被引用 2 次
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu 等AAAI 2026
- Micro-Act: Mitigate Knowledge Conflict in Question Answering via Actionable Self-ReasoningNan Huo, Jinyang Li, Bowen Qin, Ge Qu 等ACL 2025
