Emerging Data Practices: Data Work in the Era of Large Language Models
Adriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-Villacres
Abstract
Data is one of the foundational aspects of making Artificial Intelligence (AI) work as intended. As large language models (LLMs) become the epicenter of AI, it is crucial to understand better how the datasets that maintain such models are created. The emergent nature of LLMs makes it critical to understand the challenges practitioners developing Gen AI technologies face to design alternatives for better responding to Gen AI’s ethical issues. In this paper, we provide such understanding by reporting on 25 interviews with practitioners who handle data in three distinct development stages of different LLMs. Our contributions are (1) empirical evidence of how uncertainty, data practices, and reliance mechanisms change across LLMs’ development cycle; (2) how the unique qualities of LLMs impact data practices and their implications for the future of Gen AI technologies; and (3) provide three opportunities for HCI researchers interested in supporting practitioners developing Gen AI technologies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0e2b84e8-8f3b-4978-8d5e-2ab26ed1208dCited by top-tier papers3
- Digitizing the Dead: Understanding 'Data Hauntings' to Support Equitable Data Work with Human RemainsValeria Borsotti, Adam Bencard, Luigina CiolfiCHI 2026 · 1 citation
- Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to EvaluationAdriana Alvarado Garcia, Ruyuan Wan, Ozioma Collins Oguine, Karla Badillo-UrquiolaCHI 2026 · 1 citation
- From Reflection to Repair: A Scoping Review of Dataset Documentation ToolsPedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia, Heila PrecelCHI 2026 · 1 citation
Related papers
- 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research PracticesShivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li et al.CSCW 2025 · 31 citations
- Simulacrum of Stories: Examining Large Language Models as Qualitative Research ParticipantsShivani Kapania, William Agnew, Motahhare Eslami, Hoda Heidari et al.CHI 2025 · 59 citations
- Privacy Control in Conversational LLM Platforms: A Walkthrough StudyZhuoyang Li, Yanlai Wu, Yao Li, Xinning Gui et al.CHI 2026 · 1 citation
- The Impossibility of Fair LLMsJacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller et al.ACL 2025
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
