The Fault in our Stats
Alexi Turcotte, Neev Nirav Mehta
摘要
Data analysts need to be careful when they apply statistical inference techniques to data, as misuse of statistical inference methods can lead an analyst to draw the wrong conclusions. They need to be careful because, in the general case, misuse of statistics does not result in obvious problems; the numbers returned often look reasonable, and programs with misuses of statistics do not crash. In this work, we propose a technique to quickly and statically check data science programs for compliance with statistics best practice rules, including checking all assumptions made by statistical methods, as well as correcting for the multiple comparison problem, or "data dredging". This technique is predicated on a novel statistics intermediate representation, called SIR, that encodes the details most salient to statistics. We implement this technique in a tool called stat-lint, the first statistics linter, and evaluate stat-lint on 90 Python data science notebooks, finding that only 14 fully check all obligations, only two apply any correction for multiple comparisons, none validate model residuals, and over two thirds of obligations go unchecked.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Expressing and Checking Statistical AssumptionsAlexi Turcotte, Zheyuan WuFSE 2025 · 被引用 1 次
- DyLin: A Dynamic Linter for PythonAryaz Eghbali, Felix Burk, Michael PradelFSE 2025 · 被引用 1 次
- Statistical Test for Feature Selection Pipelines by Selective InferenceTomohiro Shiraishi, Tatsuya Matsukawa, Shuichi Nishino, Ichiro TakeuchiICML 2025
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 被引用 25 次
- Towards Understanding Fine-Grained Programming Mistakes and Fixing Patterns in Data ScienceWei-Hao Chen, Jia Lin Cheoh, Manthan Keim, Sabine Brunswicker 等FSE 2025 · 被引用 1 次
