Emerging Data Practices: Data Work in the Era of Large Language Models
Adriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-Villacres
摘要
Data is one of the foundational aspects of making Artificial Intelligence (AI) work as intended. As large language models (LLMs) become the epicenter of AI, it is crucial to understand better how the datasets that maintain such models are created. The emergent nature of LLMs makes it critical to understand the challenges practitioners developing Gen AI technologies face to design alternatives for better responding to Gen AI’s ethical issues. In this paper, we provide such understanding by reporting on 25 interviews with practitioners who handle data in three distinct development stages of different LLMs. Our contributions are (1) empirical evidence of how uncertainty, data practices, and reliance mechanisms change across LLMs’ development cycle; (2) how the unique qualities of LLMs impact data practices and their implications for the future of Gen AI technologies; and (3) provide three opportunities for HCI researchers interested in supporting practitioners developing Gen AI technologies.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Digitizing the Dead: Understanding 'Data Hauntings' to Support Equitable Data Work with Human RemainsValeria Borsotti, Adam Bencard, Luigina CiolfiCHI 2026 · 被引用 1 次
- Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to EvaluationAdriana Alvarado Garcia, Ruyuan Wan, Ozioma Collins Oguine, Karla Badillo-UrquiolaCHI 2026 · 被引用 1 次
- From Reflection to Repair: A Scoping Review of Dataset Documentation ToolsPedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia, Heila PrecelCHI 2026 · 被引用 1 次
相关 Paper
- 'I'm Categorizing LLM as a Productivity Tool': Examining Ethics of LLM Use in HCI Research PracticesShivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi Li 等CSCW 2025 · 被引用 31 次
- Simulacrum of Stories: Examining Large Language Models as Qualitative Research ParticipantsShivani Kapania, William Agnew, Motahhare Eslami, Hoda Heidari 等CHI 2025 · 被引用 59 次
- Privacy Control in Conversational LLM Platforms: A Walkthrough StudyZhuoyang Li, Yanlai Wu, Yao Li, Xinning Gui 等CHI 2026 · 被引用 1 次
- The Impossibility of Fair LLMsJacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller 等ACL 2025
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas 等CHI 2025 · 被引用 51 次
