Evaluating Search System Explainability with Psychometrics and Crowdsourcing
Catherine Chen, Carsten Eickhoff
摘要
As information retrieval (IR) systems, such as search engines and conversational agents, become ubiquitous in various domains, the need for transparent and explainable systems grows to ensure accountability, fairness, and unbiased results. Despite recent advances in explainable AI and IR techniques, there is no consensus on the definition of explainability. Existing approaches often treat it as a singular notion, disregarding the multidimensional definition postulated in the literature. In this paper, we use psychometrics and crowdsourcing to identify human-centered factors of explainability in Web search systems and introduce SSE (Search System Explainability), an evaluation metric for explainable IR (XIR) search systems. In a crowdsourced user study, we demonstrate SSE's ability to distinguish between explainable and non-explainable systems, showing that systems with higher scores indeed indicate greater interpretability. We hope that aside from these concrete contributions to XIR, this line of work will serve as a blueprint for similar explainability evaluation efforts in other domains of machine learning and natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- The Impact of More Transparent Interfaces on Behavior in Personalized RecommendationTobias Schnabel, Saleema Amershi, Paul N. Bennett, Peter Bailey 等SIGIR 2020 · 被引用 26 次
- Towards Explainable Search Results: A Listwise Explanation GeneratorPuxuan Yu, Razieh Rahimi, James AllanSIGIR 2022 · 被引用 26 次
- Not All Relevance Scores are Equal: Efficient Uncertainty and Calibration Modeling for Deep Retrieval ModelsDaniel Cohen, Bhaskar Mitra, Oleg Lesota, Navid Rekabsaz 等SIGIR 2021 · 被引用 18 次
相关 Paper
- Explainability for Transparent Conversational Information-SeekingWeronika Lajewska, Damiano Spina, Johanne R. Trippas, Krisztian BalogSIGIR 2024 · 被引用 17 次
- What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability MethodsJulien Colin, Thomas Fel, Rémi Cadène, Thomas SerreNeurIPS 2022 · 被引用 147 次
- EARN Fairness: Explaining, Asking, Reviewing, and Negotiating Artificial Intelligence Fairness Metrics Among StakeholdersLin Luo, Yuri Nakao, Mathieu Chollet, Hiroya Inakoshi 等CSCW 2025 · 被引用 4 次
- Dissecting users' needs for search result explanationsPrerna Juneja, Wenjuan Zhang, Alison Marie Smith-Renner, Hemank Lamba 等CHI 2024 · 被引用 4 次
- Modeling Disclosive Transparency in NLP Application DescriptionsMichael Saxon, Sharon Levy, Xinyi Wang, Alon Albalak 等EMNLP 2021 · 被引用 3 次
