What Evidence Do Language Models Find Convincing?
Alexander Wan, Eric Wallace, Dan Klein
摘要
Retrieval-augmented language models are being increasingly tasked with subjective, contentious, and conflicting queries such as "is aspartame linked to cancer". To resolve these ambiguous queries, one must search through a large range of websites and consider which, if any, of this evidence do I find convincing? In this work, we study how LLMs answer this question. In particular, we construct CON-FLICTINGQA, a dataset that pairs controversial queries with a series of real-world evidence documents that contain different facts (e.g., quantitative results), argument styles (e.g., appeals to authority), and answers (Yes or No). We use this dataset to perform sensitivity and counterfactual analyses to explore which text features most affect LLM predictions. Overall, we find that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone. Taken together, these results highlight the importance of RAG corpus quality (e.g., the need to filter misinformation), and possibly even a shift in how LLMs are trained to better align with human judgements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language ModelsThinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen 等ICLR 2026 · 被引用 69 次
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess 等ICLR 2026 · 被引用 9 次
- NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External KnowledgeHanyu Zhu, Lance Fiondella, Jiawei Yuan, Kai Zeng 等NeurIPS 2025 · 被引用 8 次
- CUB: Benchmarking Context Utilisation Techniques for Language ModelsLovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee 等ACL 2026 · 被引用 5 次
- Not All Contexts Are Equal: Teaching LLMs Credibility-aware GenerationRuotong Pan, Boxi Cao, Hongyu Lin, Xianpei Han 等EMNLP 2024 · 被引用 4 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin 等SIGIR 2021 · 被引用 297 次
相关 Paper
- LLMs Trust Humans More, That's a Problem! Unveiling and Mitigating the Authority Bias in Retrieval-Augmented GenerationYuxuan Li, Xinwei Guo, Jiashi Gao, Guanhua Chen 等ACL 2025
- Benchmarking LLM's Capability in Reasoning over Conflicting Web ReferencesYizhen Yuan, Rui Kong, Dongze Li, Yuanchun Li 等ACL 2026
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 被引用 1 次
- Conflict-Aware RAG: Multi-Stage Learning with Conflict Signals for Robust Retrieval-Augmented GenerationHaiyan Wu, Chenchen Wang, Chaoqun Sun, Chengxiong Lu 等WWW 2026
- Search Arena: Analyzing Search-Augmented LLMsMihran Miroyan, Tsung-Han Wu, Logan King, Tianle Li 等ICLR 2026 · 被引用 32 次
