JupyterLab in Retrograde: Contextual Notifications That Highlight Fairness and Bias Issues for Data Scientists
Galen Harrison, Kevin Bryson, Ahmad Emmanuel Balla Bamba, Luca Dovichi, Aleksander Herrmann Binion, Arthur Borem, Blase Ur
Abstract
Current algorithmic fairness tools focus on auditing completed models, neglecting the potential downstream impacts of iterative decisions about cleaning data and training machine learning models. In response, we developed Retrograde, a JupyterLab environment extension for Python that generates real-time, contextual notifications for data scientists about decisions they are making regarding protected classes, proxy variables, missing data, and demographic differences in model performance. Our novel framework uses automated code analysis to trace data provenance in JupyterLab, enabling these notifications. In a between-subjects online experiment, 51 data scientists constructed loan-decision models with Retrograde providing notifications continuously throughout the process, only at the end, or never. Retrograde's notifications successfully nudged participants to account for missing data, avoid using protected classes as predictors, minimize demographic differences in model performance, and exhibit healthy skepticism about their models.
• Human-centered computing → Empirical studies in HCI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 633867a7-5732-4f7d-9ff2-e538b8bb2c19Cited by top-tier papers3
- The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge WorkersHao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos et al.CHI 2025 · 690 citations
- ProactiveVA: Proactive Visual Analytics with LLM-Based UI AgentYuheng Zhao, Xueli Shu, Liwen Fan, Lin Gao et al.IEEE VIS 2025 · 4 citations
- Preventing Harmful Data Practices by using Participatory Input to Navigate the Machine Learning MultiverseJan Simson, Fiona Draxler, Samuel Mehr, Christoph KernCHI 2025 · 3 citations
Builds on15
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 428 citations
- What's Wrong with Computational Notebooks? Pain Points, Needs, and Design OpportunitiesSouti Chattopadhyay, Ishita Prasad, Austin Z. Henley, Anita Sarma et al.CHI 2020 · 162 citations
- The Landscape and Gaps in Open Source Fairness ToolkitsMichelle Seng Ah Lee, Jatinder SinghCHI 2021 · 117 citations
- Understanding and Visualizing Data Iteration in Machine LearningFred Hohman, Kanit Wongsuphasawat, Mary Beth Kery, Kayur PatelCHI 2020 · 114 citations
Related papers
- Evaluating Behavior Change Interventions for Responsible Data ScienceZiwei Dong, Keke Wu, Leilani Battle, Emily WallCHI 2026 · 1 citation
- Dead or Alive: Continuous Data Profiling for Interactive Data ScienceWill Epperson, Vaishnavi Gorantla, Dominik Moritz, Adam PererIEEE VIS 2023 · 28 citations
- Capturing and querying fine-grained provenance of preprocessing pipelines in data scienceAdriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo TorloneVLDB 2021 · 39 citations
- Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data ScienceScott Allen Cambo, Darren GergleCHI 2022 · 49 citations
- FairSense: Long-Term Fairness Analysis of ML-Enabled SystemsYining She, Sumon Biswas, Christian Kästner, Eunsuk KangICSE 2025 · 4 citations
