Expressing and Checking Statistical Assumptions
Alexi Turcotte, Zheyuan Wu
Abstract
Literate programming environments like Jupyter and R Markdown notebooks, coupled with easy-to-use languages like Python and R, put a plethora of statistical methods right at a data analyst's fingertips. But are these methods being used correctly? Statistical methods make statistical assumptions about samples being analyzed, and in many cases produce reasonable looking results even if assumptions are not met.
We propose an approach that allows library developers to annotate functions with statistical assumptions, phrases them as hypotheses about the data, and inserts hypothesis tests investigating the likelihood that the assumption is met; this way, analysts using these functions will have their data checked automatically. We implement this approach in two tools: prob-check-py for Python, and prob-check-r for R, and to evaluate them we identify common hypothesis testing and statistical modeling functions, annotate them with the relevant statistical assumptions, and run 128 Kaggle notebooks that use those methods to identify misuses. Our investigation reveals statistically significant evidence against assumptions in 84.38% of surveyed notebooks, and in 53.36% of calls to annotated functions. In the case of hypothesis tests, had an equivalent test that did not make these assumptions been chosen, a different conclusion would have been drawn in 11.51% of cases.
CCS Concepts: • Software and its engineering → General programming languages; • Mathematics of computing → Probability and statistics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Typilus: neural type hintsMiltiadis Allamanis, Earl T. Barr, Soline Ducousso, Zheng GaoPLDI 2020 · 92 citations
- Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data RelationshipsEunice Jun, Audrey Seo, Jeffrey Heer, René JustCHI 2022 · 24 citations
- Designing types for R, empiricallyAlexi Turcotte, Aviral Goel, Filip Krikava, Jan VitekOOPSLA 2020 · 9 citations
- Certifying Certainty and Uncertainty in Approximate Membership Query StructuresKiran Gopinathan, Ilya SergeyCAV 2020 · 9 citations
- Does blame shifting work?Lukas Lazarek, Alexis King, Samanvitha Sundar, Robert Bruce Findler et al.POPL 2020 · 7 citations
Related papers
- The Fault in our StatsAlexi Turcotte, Neev Nirav MehtaASE 2025
- Language-Agnostic Static Analysis of Probabilistic ProgramsMarkus Böck, Michael Schröder, Jürgen CitoASE 2024 · 4 citations
- Towards Understanding Fine-Grained Programming Mistakes and Fixing Patterns in Data ScienceWei-Hao Chen, Jia Lin Cheoh, Manthan Keim, Sabine Brunswicker et al.FSE 2025 · 1 citation
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 25 citations
- The raise of machine learning hyperparameter constraints in Python codeIngkarat Rak-amnouykit, Ana L. Milanova, Guillaume Baudart, Martin Hirzel et al.ISSTA 2022 · 1 citation
