The Craft and Coordination of Data Curation: Complicating Workflow Views of Data Science
Andrea K. Thomer, Dharma Akmon, Jeremy York, Allison R. B. Tyler, Faye Polasek, Sara Lafia, Libby Hemphill, Elizabeth Yakel
摘要
Data curation is the process of making a dataset fit-for-use and archivable. It is critical to data-intensive science because it makes complex data pipelines possible, studies reproducible, and data reusable. Yet the complexities of the hands-on, technical, and intellectual work of data curation is frequently overlooked or downplayed. Obscuring the work of data curation not only renders the labor and contributions of data curators invisible but also hides the impact that curators' work has on the later usability, reliability, and reproducibility of data. To better understand the work and impact of data curation, we conducted a close examination of data curation at a large social science data repository, the Inter-university Consortium for Political and Social Research (ICPSR). We asked: What does curatorial work entail at ICPSR, and what work is more or less visible to different stakeholders and in different contexts? And, how is that curatorial work coordinated across the organization? We triangulated accounts of data curation from interviews and records of curation in Jira tickets to develop a rich and detailed account of curatorial work. While we identified numerous curatorial actions performed by ICPSR curators, we also found that curators rely on a number of craft practices to perform their jobs. The reality of their work practices defies the rote sequence of events implied by many life cycle or workflow models. Further, we show that craft practices are needed to enact data curation best practices and standards. The craft that goes into data curation is often invisible to end users, but it is well recognized by ICPSR curators and their supervisors. Explicitly acknowledging and supporting data curators as craftspeople is important in creating sustainable and successful curatorial infrastructures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Organizing Oceanographic Infrastructure: The Work of Making a Software Pipeline RepurposableAndrew B. Neang, Will Sutherland, David Ribes, Charlotte P. LeeCSCW 2023 · 被引用 7 次
- OpenForge: Probabilistic Metadata IntegrationTianji Cong, Fatemeh Nargesian, Junjie Xing, H. V. JagadishVLDB 2025 · 被引用 1 次
它引用的顶会 Paper5
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 被引用 169 次
- Data Integration as Coordination: The Articulation of Data Work in an Ocean Science CollaborationAndrew B. Neang, Will Sutherland, Michael W. Beach, Charlotte P. LeeCSCW 2020 · 被引用 74 次
- Orienting, Framing, Bridging, Magic, and Counseling: How Data Scientists Navigate the Outer Loop of Client Collaborations in Industry and AcademiaSean Kross, Philip J. GuoCSCW 2021 · 被引用 35 次
- 'Yes, I comply!': Motivations and Practices around Research Data Management and Reuse across Scientific FieldsSebastian S. Feger, Pawel W. Wozniak, Lars Lischke, Albrecht SchmidtCSCW 2020 · 被引用 29 次
相关 Paper
- The New Reality of Reproducibility: The Role of Data Work in Scientific ResearchMelanie Feinberg, Will Sutherland, Sarah Beth Nelson, Mohammad Hossein Jarrahi 等CSCW 2020 · 被引用 21 次
- Unveiling Practices of Customer Service Content Curators of Conversational AgentsHeloisa Candello, Claudio S. Pinhanez, Michael J. Muller, Mairieli WesselCSCW 2022 · 被引用 8 次
- The Art and Practice of Data Science Pipelines: A Comprehensive Study of Data Science Pipelines In Theory, In-The-Small, and In-The-LargeSumon Biswas, Mohammad Wardat, Hridesh RajanICSE 2022 · 被引用 64 次
- How Domain Experts Work with Data: Situating Data Science in the Practices and Settings of CraftworkJu-Yeon Jung, Tom Steinberger, John L. King, Mark S. AckermanCSCW 2022 · 被引用 24 次
- How Researchers De-Identify Data in PracticeWentao Guo, Paige Pepitone, Adam J. Aviv, Michelle L. MazurekUSENIX Security 2025
