Treading the Transparency Tightrope: A Taxonomy of Risks and Benefits of Foundation Model Data Transparency for Transparency Advocates
Morgan Klaus Scheuerman, Wiebke Hutiri, Aida Rahmattalabi, Victoria Matthews, Alice Xiang, Jerone Theodore Alexander Andrews
Abstract
Data powering AI is often opaque. Researchers, NGOs, and law and policy leaders have called for greater transparency about how data is used for training, fine-tuning, and evaluation. While data transparency is often championed as crucial, what it concretely enables is largely implicit. Similarly, the concerns developers seem to have about transparency go unstated. This lack of clarity has led some researchers to critique transparency demands as disconnected from the actual benefits—or risks—to specific stakeholders. We analyze documentation from four stakeholder groups to create a taxonomy of the risks and benefits of dataset transparency. Data transparency is perceived as either a risk or a benefit given a stakeholder’s position, rather than wholesale. We also propose data availability and data documentation as two lenses through which to consider transparency. We discuss how best to strategically promote situational data transparency that takes into account the relationship between stakeholder position, transparency modality, and benefits/risks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 20ffc8f7-7229-438b-ac81-c69b249cea25Related papers
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and DesiderataAmy Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna M. Wallach et al.CSCW 2022 · 58 citations
- "It just requires so much more creativity": Barriers and Workarounds to Gathering Information for AI ContestationSohini Upadhyay, Dasha Pruss, Alicia DeVrio, Krzysztof Z. Gajos et al.CHI 2026 · 1 citation
- Regulating AI: Where U.S. State Policy and HCI (Mis)alignNino Migineishvili, Alice Gao, Adinawa Adjagbodjou, Dhanaraj Thakur et al.CHI 2026 · 1 citation
- Contributing to Accessibility Datasets: Reflections on Sharing Study Data by Blind PeopleRie Kamikubo, Kyungjun Lee, Hernisa KacorriCHI 2023 · 16 citations
- "You Can either Blame Technology or Blame a Person..." - A Conceptual Model of Users' AI-Risk Perception as a Tool for HCILena Recki, Dennis Lawo, Veronika Krauß, Dominik Pins et al.CSCW 2024 · 6 citations
