Evaluating the Interpretability of Generative Models by Interactive Reconstruction
Andrew Slavin Ross, Nina Chen, Elisa Zhao Hang, Elena L. Glassman, Finale Doshi-Velez
Abstract
For machine learning models to be most useful in numerous sociotechnical systems, many have argued that they must be human-interpretable. However, despite increasing interest in interpretability, there remains no firm consensus on how to measure it. This is especially true in representation learning, where interpretability research has focused on “disentanglement” measures only applicable to synthetic datasets and not grounded in human factors. We introduce a task to quantify the human-interpretability of generative model representations, where users interactively modify representations to reconstruct target instances. On synthetic datasets, we find performance on this task much more reliably differentiates entangled and disentangled models than baseline approaches. On a real dataset, we find it differentiates between representation learning methods widely believed but never shown to produce more or less interpretable models. In both cases, we ran small-scale think-aloud studies and large-scale experiments on Amazon Mechanical Turk to confirm that our qualitative and quantitative results agreed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- DirectGPT: A Direct Manipulation Interface to Interact with Large Language ModelsDamien Masson, Sylvain Malacria, Géry Casiez, Daniel VogelCHI 2024 · 104 citations
- Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language ModelsTae Soo Kim, Yoonjoo Lee, Minsuk Chang, Juho KimUIST 2023 · 55 citations
- On Selective, Mutable and Dialogic XAI: a Review of What Users Say about Different Types of Interactive ExplanationsAstrid Bertrand, Tiphaine Viard, Rafik Belloum, James R. Eagan et al.CHI 2023 · 53 citations
- GANSlider: How Users Control Generative Models for Images using Multiple Sliders with and without Feedforward InformationHai Dang, Lukas Mecke, Daniel BuschekCHI 2022 · 38 citations
- Understanding Instance-based Interpretability of Variational Auto-EncodersZhifeng Kong, Kamalika ChaudhuriNeurIPS 2021 · 32 citations
Builds on4
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana et al.CHI 2020 · 541 citations
- COGAM: Measuring and Moderating Cognitive Load in Machine Learning Model ExplanationsAshraf M. Abdul, Christian von der Weth, Mohan S. Kankanhalli, Brian Y. LimCHI 2020 · 92 citations
- A Loss Function for Generative Neural Networks Based on Watson's Perceptual ModelSteffen Czolbe, Oswin Krause, Ingemar J. Cox, Christian IgelNeurIPS 2020 · 71 citations
Related papers
- Evaluating the Disentanglement of Deep Generative Models through Manifold TopologySharon Zhou, Eric Zelikman, Fred Lu, Andrew Y. Ng et al.ICLR 2021 · 29 citations
- Towards Robust Metrics for Concept Representation EvaluationMateo Espinosa Zarlenga, Pietro Barbiero, Zohreh Shams, Dmitry Kazhdan et al.AAAI 2023 · 32 citations
- Transferring disentangled representations: bridging the gap between synthetic and real imagesJacopo Dapueto, Nicoletta Noceti, Francesca OdoneNeurIPS 2024 · 3 citations
- Where and What? Examining Interpretable Disentangled RepresentationsXinqi Zhu, Chang Xu, Dacheng TaoCVPR 2021
- Theory and Evaluation Metrics for Learning Disentangled RepresentationsKien Do, Truyen TranICLR 2020 · 107 citations
