The Viability of Crowdsourcing for RAG Evaluation
Lukas Gienapp, Tim Hagen, Maik Fröbe, Matthias Hagen, Benno Stein, Martin Potthast, Harrisen Scells
摘要
How good are humans at writing and judging responses in retrievalaugmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. We present the Crowd RAG Corpus 2025 (CrowdRAG-25), which consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC RAG'24 track, across the three discourse styles 'bulleted list', 'essay', and 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. Human pairwise judgments provide reliable and cost-effective results compared to LLM-based pairwise or human/LLM-based pointwise judgments, as well as automated comparisons with human-written reference responses. All our data and tools are freely available. 1
• Information systems → Evaluation of retrieval results; Language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 被引用 531 次
- Large Language Models can Accurately Predict Searcher PreferencesPaul Thomas, Seth Spielman, Nick Craswell, Bhaskar MitraSIGIR 2024 · 被引用 153 次
- Human Feedback is not Gold StandardTom Hosking, Phil Blunsom, Max BartoloICLR 2024 · 被引用 96 次
相关 Paper
- MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented GenerationMaría Andrea Cruz Blandón, Jayasimha Talur, Bruno Charron, Dong Liu 等ACL 2025
- A Comparison of Conversational Models and Humans in Answering Technical Questions: the Firefox CaseJoão Correia, Daniel Coutinho, Marco Castelluccio, Caio Barbosa 等ICSE 2026
- MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval BenchmarksJunhao Ruan, Abudukeyumu Abudula, Bei Li, Yongjing Yin 等ACL 2026
- Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam GenerationGauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, Laurent CallotICML 2024 · 被引用 35 次
- Are Large Language Models Good at Utility Judgments?Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等SIGIR 2024 · 被引用 20 次
