Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
Luke Guerdan, Devansh Saxena, Stevie Chancellor, Zhiwei Steven Wu, Kenneth Holstein
摘要
Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the ''authenticity'' of student writing or the ''healthcare need'' of a patient. Yet the process by which data scientists translate fuzzy concepts into a concrete, proxy target variable remains poorly understood. We interview fifteen data scientists in education (N=8) and healthcare (N=7) to understand how they construct target variables for predictive modeling tasks. Our findings suggest that data scientists construct target variables through a bricolage process, in which they use creative and pragmatic approaches to make do with the limited data at hand. Data scientists attempt to satisfy five major criteria for a target variable through bricolage: validity, simplicity, predictability, portability, and resource requirements. To achieve this, data scientists adaptively apply problem (re)formulation strategies, such as swapping out one candidate target variable for another when the first fails to meet certain criteria (e.g., predictability), or composing multiple outcomes into a single target variable to capture a more holistic set of modeling objectives. Based on our findings, we present opportunities for future HCI, CSCW, and ML research to better support the art and science of target variable construction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran 等UIST 2024 · 被引用 143 次
- Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision SupportAnna Kawakami, Venkatesh Sivaraman, Hao Fei Cheng, Logan Stapleton 等CHI 2022 · 被引用 137 次
- Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and ProcessNadia Nahar, Shurui Zhou, Grace A. Lewis, Christian KästnerICSE 2022 · 被引用 122 次
相关 Paper
- DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationArpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi 等CHI 2023 · 被引用 13 次
- How Domain Experts Work with Data: Situating Data Science in the Practices and Settings of CraftworkJu-Yeon Jung, Tom Steinberger, John L. King, Mark S. AckermanCSCW 2022 · 被引用 24 次
- Rethinking Dataset Discovery with DataScoutRachel Lin, Bhavya Chopra, Wenjing Lin, Shreya Shankar 等UIST 2025 · 被引用 2 次
- Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling TasksVenkatesh Sivaraman, Anika Vaishampayan, Xiaotong Li, Brian R. Buck 等CHI 2025 · 被引用 1 次
- The Work of Infrastructural Bricoleurs in Building Civic Data DashboardsFiraz Peer, Carl F. DiSalvoCSCW 2022 · 被引用 15 次
