Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance Judgments
Kevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca Demartini
摘要
In Information Retrieval (IR) evaluation, preference judgments are collected by presenting to the assessors a pair of documents and asking them to select which of the two, if any, is the most relevant. This is an alternative to the classic relevance judgment approach, in which human assessors judge the relevance of a single document on a scale; such an alternative allows to make relative rather than absolute judgments of relevance. While preference judgments are easier for human assessors to perform, the number of possible document pairs to be judged is usually so high that it makes it unfeasible to judge them all. Thus, following a similar idea to pooling strategies for single document relevance judgments where the goal is to sample the most useful documents to be judged, in this work we focus on analyzing alternative ways to sample document pairs to judge, in order to maximize the value of a fixed number of preference judgments that can feasibly be collected. Such value is defined as how well we can evaluate IR systems given a budget, that is, a fixed number of human preference judgments that may be collected. By relying on several datasets featuring relevance judgments gathered by means of experts and crowdsourcing, we experimentally compare alternative strategies to select document pairs and show how different strategies lead to different IR evaluation result quality levels. Our results show that, by using the appropriate procedure, it is possible to achieve good IR evaluation results with a limited number of preference judgments, thus confirming the feasibility of using preference judgments to create IR evaluation collections.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Extending Label Aggregation Models with a Gaussian Process to Denoise Crowdsourcing LabelsDan Li, Maarten de RijkeSIGIR 2023 · 被引用 21 次
- Human Preferences as Dueling BanditsXinyi Yan, Chengxi Luo, Charles L. A. Clarke, Nick Craswell 等SIGIR 2022 · 被引用 9 次
相关 Paper
- Good Evaluation Measures based on Document PreferencesTetsuya Sakai, Zhaohao ZengSIGIR 2020 · 被引用 14 次
- Hear Me Out: A Study on the Use of the Voice Modality for Crowdsourced Relevance AssessmentsNirmal Roy, Agathe Balayn, David Maxwell, Claudia HauffSIGIR 2023 · 被引用 1 次
- Hybrid Pooling with LLMs via Relevance Context LearningDavid Otero, Javier ParaparSIGIR 2026
- Active Ordinal Querying for Tuplewise Similarity LearningGregory Canal, Stefano Fenu, Christopher RozellAAAI 2020 · 被引用 10 次
- Evaluation Measures Based on Preference GraphsCharles L. A. Clarke, Chengxi Luo, Mark D. SmuckerSIGIR 2021 · 被引用 5 次
