Sensemaking in User-Driven Algorithm Auditing: A Case Study on Gender Bias in an Image Captioning Model
Behnoosh Mohammadzadeh, Jules Françoise, Michèle Gouiffès, Baptiste Caramiaux
Abstract
Non-experts increasingly engage in user-driven algorithm auditing, interacting directly with AI systems to probe, document, and reflect on biased behavior. Yet, auditing remains challenging due to model opacity and limited support for navigating and interpreting outputs. This paper explores the design and evaluation of interfaces grounded in the sensemaking framework to support non-experts in auditing gender bias in image captioning. In a between-subjects study, 60 participants audited an image captioning model using one of three interface conditions: a Baseline interface, a Masking Tool for image manipulation, or a Filtering Tool for organizing captions. Our findings show that interface design shaped what participants noticed, how they interpreted model behavior, and supported their hypotheses. The Image Masking Tool enabled fine-grained testing of visual cues and context, while the Text Filtering Tool revealed broader asymmetries in gendered language. We argue that incorporating sensemaking into auditing practices can advance accountability and transparency in machine learning systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 229ea879-adf3-4941-aeca-912fee308541Builds on17
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Problematic Machine Behavior: A Systematic Literature Review of Algorithm AuditsJack BandyCSCW 2021 · 190 citations
- Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic BehaviorsHong Shen, Alicia DeVos, Motahhare Eslami, Kenneth HolsteinCSCW 2021 · 156 citations
- Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for SupportMichael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan et al.CSCW 2022 · 149 citations
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language ModelsSangho Suh, Bryan Min, Srishti Palani, Haijun XiaUIST 2023 · 147 citations
Related papers
- Vipera: Blending Visual and LLM-Driven Guidance for Systematic Auditing of Text-to-Image Generative AIYanwei Huang, Wesley Hanwen Deng, Sijia Xiao, Motahhare Eslami et al.CHI 2026 · 2 citations
- Toward User-Driven Algorithm Auditing: Investigating users' strategies for uncovering harmful algorithmic behaviorAlicia DeVos, Aditi Dhabalia, Hong Shen, Kenneth Holstein et al.CHI 2022 · 96 citations
- Learning AI Auditing: A Case Study of Teenagers Auditing a Generative AI ModelLuis Morales-Navarro, Michelle A. Gan, Evelyn Yu, Lauren Vogelstein et al.CSCW 2025 · 6 citations
- Interpretable Debiasing of Vision-Language Models for Social FairnessNa Min An, Yoonna Jang, Yusuke Hirota, Ryo Hachiuma et al.CVPR 2026 · 7 citations
- To "See" is to Stereotype: Image Tagging Algorithms, Gender Recognition, and the Accuracy-Fairness Trade-offPinar Barlas, Kyriakos Kyriakou, Olivia Guest, Styliani Kleanthous et al.CSCW 2020 · 40 citations
