Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
Luke Guerdan, Devansh Saxena, Stevie Chancellor, Zhiwei Steven Wu, Kenneth Holstein
Abstract
Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the ''authenticity'' of student writing or the ''healthcare need'' of a patient. Yet the process by which data scientists translate fuzzy concepts into a concrete, proxy target variable remains poorly understood. We interview fifteen data scientists in education (N=8) and healthcare (N=7) to understand how they construct target variables for predictive modeling tasks. Our findings suggest that data scientists construct target variables through a bricolage process, in which they use creative and pragmatic approaches to make do with the limited data at hand. Data scientists attempt to satisfy five major criteria for a target variable through bricolage: validity, simplicity, predictability, portability, and resource requirements. To achieve this, data scientists adaptively apply problem (re)formulation strategies, such as swapping out one candidate target variable for another when the first fails to meet certain criteria (e.g., predictability), or composing multiple outcomes into a single target variable to capture a more holistic set of modeling objectives. Based on our findings, we present opportunities for future HCI, CSCW, and ML research to better support the art and science of target variable construction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea91c90d-3158-4685-ba84-e25049bd8b19Cited by top-tier papers1
Ask how each one uses itBuilds on26
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 260 citations
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human PreferencesShreya Shankar, J. D. Zamfirescu-Pereira, Bjoern Hartmann, Aditya G. Parameswaran et al.UIST 2024 · 143 citations
- Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision SupportAnna Kawakami, Venkatesh Sivaraman, Hao Fei Cheng, Logan Stapleton et al.CHI 2022 · 137 citations
- Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and ProcessNadia Nahar, Shurui Zhou, Grace A. Lewis, Christian KästnerICSE 2022 · 122 citations
Related papers
- DataPilot: Utilizing Quality and Usage Information for Subset Selection during Visual Data PreparationArpit Narechania, Fan Du, Atanu R. Sinha, Ryan A. Rossi et al.CHI 2023 · 13 citations
- How Domain Experts Work with Data: Situating Data Science in the Practices and Settings of CraftworkJu-Yeon Jung, Tom Steinberger, John L. King, Mark S. AckermanCSCW 2022 · 24 citations
- Rethinking Dataset Discovery with DataScoutRachel Lin, Bhavya Chopra, Wenjing Lin, Shreya Shankar et al.UIST 2025 · 2 citations
- Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling TasksVenkatesh Sivaraman, Anika Vaishampayan, Xiaotong Li, Brian R. Buck et al.CHI 2025 · 1 citation
- The Work of Infrastructural Bricoleurs in Building Civic Data DashboardsFiraz Peer, Carl F. DiSalvoCSCW 2022 · 15 citations
