Algorithmic Vibe in Information Retrieval
Ali Montazeralghaem, Nick Craswell, Ryen W. White, Ahmed Hassan Awadallah, Byungki Byun
Abstract
When information retrieval systems return a ranked list of results in response to a query, they may be choosing from a large set of candidate results that are equally useful and relevant. This means we might be able to identify a difference between rankers A and B, where ranker A systematically prefers a certain type of relevant results. Ranker A may have this systematic difference (different “vibe”) without having systematically better or worse results according to standard information retrieval metrics. We first show that a vibe difference can exist, comparing two publicly available rankers, where the one that is trained on health-related queries will systematically prefer health-related results, even for non-health queries. We define a vibe metric that lets us see the words that a ranker prefers. We investigate the vibe of search engine clicks vs. human labels. We perform an initial study into correcting for vibe differences to make ranker A more like ranker B via changes in negative sampling during training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b7af9c0-d225-4b96-94fd-9f91bea2605fBuilds on3
- Societal Biases in Retrieved Contents: Measurement Framework and Adversarial Mitigation of BERT RankersNavid Rekabsaz, Simone Kopeinik, Markus SchedlSIGIR 2021 · 55 citations
- A Reinforcement Learning Framework for Relevance FeedbackAli Montazeralghaem, Hamed Zamani, James AllanSIGIR 2020 · 38 citations
- Unsupervised Corpus Aware Language Model Pre-training for Dense Passage RetrievalLuyu Gao, Jamie CallanACL 2022
Related papers
- VibeCheck: Discover and Quantify Qualitative Differences in Large Language ModelsLisa Dunlap, Krishna Mandal, Trevor Darrell, Jacob Steinhardt et al.ICLR 2025
- Correcting for Selection Bias in Learning-to-rank SystemsZohreh Ovaisi, Ragib Ahsan, Yifan Zhang, Kathryn Vasilaky et al.WWW 2020 · 123 citations
- Models Versus Satisfaction: Towards a Better Understanding of Evaluation MetricsFan Zhang, Jiaxin Mao, Yiqun Liu, Xiaohui Xie et al.SIGIR 2020 · 35 citations
- Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance JudgmentsKevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca DemartiniWWW 2022 · 7 citations
- Do Affective Cues Validate Behavioural Metrics for Search?Daniel McDuff, Paul Thomas, Nick Craswell, Kael Rowan et al.SIGIR 2021 · 12 citations
