Algorithmic Vibe in Information Retrieval
Ali Montazeralghaem, Nick Craswell, Ryen W. White, Ahmed Hassan Awadallah, Byungki Byun
摘要
When information retrieval systems return a ranked list of results in response to a query, they may be choosing from a large set of candidate results that are equally useful and relevant. This means we might be able to identify a difference between rankers A and B, where ranker A systematically prefers a certain type of relevant results. Ranker A may have this systematic difference (different “vibe”) without having systematically better or worse results according to standard information retrieval metrics. We first show that a vibe difference can exist, comparing two publicly available rankers, where the one that is trained on health-related queries will systematically prefer health-related results, even for non-health queries. We define a vibe metric that lets us see the words that a ranker prefers. We investigate the vibe of search engine clicks vs. human labels. We perform an initial study into correcting for vibe differences to make ranker A more like ranker B via changes in negative sampling during training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Societal Biases in Retrieved Contents: Measurement Framework and Adversarial Mitigation of BERT RankersNavid Rekabsaz, Simone Kopeinik, Markus SchedlSIGIR 2021 · 被引用 55 次
- A Reinforcement Learning Framework for Relevance FeedbackAli Montazeralghaem, Hamed Zamani, James AllanSIGIR 2020 · 被引用 38 次
- Unsupervised Corpus Aware Language Model Pre-training for Dense Passage RetrievalLuyu Gao, Jamie CallanACL 2022
相关 Paper
- VibeCheck: Discover and Quantify Qualitative Differences in Large Language ModelsLisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt 等ICLR 2025
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky 等WWW 2020 · 被引用 123 次
- Models Versus Satisfaction: Towards a Better Understanding of Evaluation MetricsFan Zhang, Jiaxin Mao, Yiqun Liu, Xiaohui Xie 等SIGIR 2020 · 被引用 35 次
- Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance JudgmentsKevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca DemartiniWWW 2022 · 被引用 7 次
- Do Affective Cues Validate Behavioural Metrics for Search?Daniel McDuff, Paul Thomas, Nick Craswell, Kael Rowan 等SIGIR 2021 · 被引用 12 次
