Lune

CHI2021Top-tier venue

Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency

Ariful Islam Anik, Andrea Bunt

2021Year
99Citations
20Top-tier citations

Abstract

Training datasets fundamentally impact the performance of machine learning systems.

Any biases introduced during training (implicit or explicit) are often reflected in the system's behaviors leading to questions about fairness and loss of trust in the system. Yet, information on training data is rarely communicated to the stakeholders. In this thesis, I explore the concept of data-centric explanations for machine learning systems that describe the training data to end-users. I design data-centric explanations that focus on providing information on training data. Through a formative study, I investigate the potential utility of such an approach and the data-centric information that users find most compelling. In a second study, I investigate reactions to the explanations across four different system scenarios. The results show that data-centric explanations can impact how users judge the trustworthiness of a system and can assist users in assessing fairness. I discuss the implications of the findings for designing explanations to support users' perception of machine learning systems.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d177a9ad-5c66-4d24-ba19-abf36c0c667a

Cited by top-tier papers20

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines