Expressing and Checking Statistical Assumptions
Alexi Turcotte, Zheyuan Wu
摘要
Literate programming environments like Jupyter and R Markdown notebooks, coupled with easy-to-use languages like Python and R, put a plethora of statistical methods right at a data analyst's fingertips. But are these methods being used correctly? Statistical methods make statistical assumptions about samples being analyzed, and in many cases produce reasonable looking results even if assumptions are not met.
We propose an approach that allows library developers to annotate functions with statistical assumptions, phrases them as hypotheses about the data, and inserts hypothesis tests investigating the likelihood that the assumption is met; this way, analysts using these functions will have their data checked automatically. We implement this approach in two tools: prob-check-py for Python, and prob-check-r for R, and to evaluate them we identify common hypothesis testing and statistical modeling functions, annotate them with the relevant statistical assumptions, and run 128 Kaggle notebooks that use those methods to identify misuses. Our investigation reveals statistically significant evidence against assumptions in 84.38% of surveyed notebooks, and in 53.36% of calls to annotated functions. In the case of hypothesis tests, had an equivalent test that did not make these assumptions been chosen, a different conclusion would have been drawn in 11.51% of cases.
CCS Concepts: • Software and its engineering → General programming languages; • Mathematics of computing → Probability and statistics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Typilus: neural type hintsMiltiadis Allamanis, Earl T. Barr, Soline Ducousso, Zheng GaoPLDI 2020 · 被引用 92 次
- Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data RelationshipsEunice Jun, Audrey Seo, Jeffrey Heer, René JustCHI 2022 · 被引用 24 次
- Designing types for R, empiricallyAlexi Turcotte, Aviral Goel, Filip Krikava, Jan VitekOOPSLA 2020 · 被引用 9 次
- Certifying Certainty and Uncertainty in Approximate Membership Query StructuresKiran Gopinathan, Ilya SergeyCAV 2020 · 被引用 9 次
- Does blame shifting work?Lukas Lazarek, Alexis King, Samanvitha Sundar, Robert Bruce Findler 等POPL 2020 · 被引用 7 次
相关 Paper
- The Fault in our StatsAlexi Turcotte, Neev Nirav MehtaASE 2025
- Language-Agnostic Static Analysis of Probabilistic ProgramsMarkus Böck, Michael Schröder, Jürgen CitoASE 2024 · 被引用 4 次
- Towards Understanding Fine-Grained Programming Mistakes and Fixing Patterns in Data ScienceWei-Hao Chen, Jia Lin Cheoh, Manthan Keim, Sabine Brunswicker 等FSE 2025 · 被引用 1 次
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 被引用 25 次
- The raise of machine learning hyperparameter constraints in Python codeIngkarat Rak-amnouykit, Ana L. Milanova, Guillaume Baudart, Martin Hirzel 等ISSTA 2022 · 被引用 1 次
