MULTIVERSE: Mining Collective Data Science Knowledge from Code on the Web to Suggest Alternative Analysis Approaches
Mike A. Merrill, Ge Zhang, Tim Althoff
摘要
Data analyses are based on a series of "decision points" including data filtering, feature operationalization and selection, model specification, and parametric assumptions. "Multiverse Analysis" research has shown that a lack of exploration of these decisions can lead to non-robust conclusions based on highly sensitive decision points. Importantly, even if myopic analyses are technically correct, analysts' focus on one set of decision points precludes them from exploring alternate formulations that may produce very different results. Prior work has also shown that analysts' exploration is often limited based on their training, domain, and personal experience. However, supporting analysts in exploring alternative approaches is challenging and typically requires expert feedback that is costly and hard to scale. Here, we formulate the tasks of identifying decision points and suggesting alternative analysis approaches as a classification task and a sequence-to-sequence prediction task, respectively. We leverage public collective data analysis knowledge in the form of code submissions to the popular data science platform Kaggle to build the first predictive model which supports Multiverse Analysis. Specifically, we mine this code repository for 70k small differences between 40k submissions, and demonstrate that these differences often highlight key decision points and alternative approaches in their respective analyses. We leverage information on relationships within libraries through neural graph representation learning in a multitask learning framework. We demonstrate that our model, MULTIVERSE, is able to correctly predict decision points with up to 0.81 ROC AUC, and alternative code snippets with up to 50.3% GLEU, and that it performs favorably compared to a suite of baselines and ablations. We show that when our model has perfect information about the location of decision points, say provided by the analyst, its performance increases significantly from 50.3% to 73.4% GLEU. Finally, we show through a human evaluation that real data analysts find alternatives provided by MULTIVERSE to be more reasonable, acceptable, and syntactically correct than alternatives from comparable baselines, including other transformer-based seq2seq models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Boba: Authoring and Visualizing Multiverse AnalysesYang Liu, Alex Kale, Tim Althoff, Jeffrey HeerIEEE VIS 2020 · 被引用 79 次
- Low-Dimensional Hyperbolic Knowledge Graph EmbeddingsInes Chami, Adva Wolf, Da-Cheng Juan, Frederic Sala 等ACL 2020 · 被引用 48 次
- Paths Explored, Paths Omitted, Paths Obscured: Decision Points & Selective Reporting in End-to-End Data AnalysisYang Liu, Tim Althoff, Jeffrey HeerCHI 2020 · 被引用 40 次
相关 Paper
- multiverse: Multiplexing Alternative Data Analyses in R NotebooksAbhraneel Sarma, Alex Kale, Michael Jongho Moon, Nathan Taback 等CHI 2023 · 被引用 22 次
- Multiverse Notebook: Shifting Data Scientists to Time TravelersShigeyuki Sato, Tomoki NakamaruOOPSLA 2024 · 被引用 3 次
- Understanding and Supporting Debugging Workflows in Multiverse AnalysisKen Gu, Eunice Jun, Tim AlthoffCHI 2023 · 被引用 11 次
- Milliways: Taming Multiverses through Principled Evaluation of Data Analysis PathsAbhraneel Sarma, Kyle Hwang, Jessica Hullman, Matthew KayCHI 2024 · 被引用 17 次
- How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz StudyKen Gu, Madeleine Grunde-McLaughlin, Andrew M. McNutt, Jeffrey Heer 等CHI 2024 · 被引用 37 次
