Statically Analyzing the Dataflow of R Programs
Florian Sihler, Matthias Tichy
Abstract
The R programming language is primarily designed for statistical computing and mostly used by researchers without a background in computer science. R provides a wide range of dynamic features and peculiarities that are difficult to analyze statically like dynamic scoping and lazy evaluation with dynamic side effects. At the same time, the R ecosystem lacks sophisticated analysis tools that support researchers in understanding and improving their code. In this paper, we present a novel static dataflow analysis framework for the R programming language that is capable of handling the dynamic nature of R programs and produces the dataflow graph of given R programs. This graph can be essential in a range of analyses, including program slicing, which we implement as a proof of concept. The core analysis works as a stateful fold over a normalized version of the abstract syntax tree of the R program, which tracks (re-)definitions, values, function calls, side effects, external files, and a dynamic control flow to produce one dataflow graph per program. We evaluate the correctness of our analysis using output equivalence testing on a manually curated dataset of 779 sensible slicing points from executable real-world R scripts. Additionally, we use a set of systematic test cases based on the capabilities of the R language and the implementation of the R interpreter and measure the runtimes well as the memory consumption on a set of 4,230 real-world R scripts and 20,815 packages available on R’s package manager CRAN. Furthermore, we evaluate the recall of our program slicer, its accuracy using shrinking, and its improvement over the state of the art. We correctly analyze almost all programs in our equivalence test suite, preserving the identical output for 99.7 % of the manually curated slicing points. On average, we require 576 ms to analyze the dataflow and around 213 kB to store the graph of a research script. This shows that our analysis is capable of analyzing real-world sources quickly and correctly. Our slicer achieves an average reduction of 84.8 % of tokens indicating its potential to improve program comprehension.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 33cdcbf9-0b7a-4cd8-8199-0d6c3e43cc4aRelated papers
- What we eval in the shadows: a large-scale study of eval in R programsAviral Goel, Pierre Donat-Bouillud, Filip Krikava, Christoph M. Kirsch et al.OOPSLA 2021 · 3 citations
- Promises are made to be broken: migrating R to strict semanticsAviral Goel, Jan Jecmen, Sebastián Krynski, Olivier Flückiger et al.OOPSLA 2021 · 2 citations
- Designing types for R, empiricallyAlexi Turcotte, Aviral Goel, Filip Krikava, Jan VitekOOPSLA 2020 · 9 citations
- Bolt-on, Compact, and Rapid Program Slicing for Notebooks [Scalable Data Science]Shreya Shankar, Stephen Macke, Sarah E. Chasins, Andrew Head et al.VLDB 2022 · 17 citations
- A Study of Undefined Behavior Across Foreign Function Boundaries in Rust LibrariesIan McCormack, Joshua Sunshine, Jonathan AldrichICSE 2025 · 5 citations
