rTisane: Externalizing conceptual models for data analysis prompts reconsideration of domain assumptions and facilitates statistical modeling
Eunice Jun, Edward Misback, Jeffrey Heer, René Just
Abstract
Statistical models should accurately reflect analysts’ domain knowledge about variables and their relationships. While recent tools let analysts express these assumptions and use them to produce a resulting statistical model, it remains unclear what analysts want to express and how externalization impacts statistical model quality. This paper addresses these gaps. We first conduct an exploratory study of analysts using a domain-specific language (DSL) to express conceptual models. We observe a preference for detailing how variables relate and a desire to allow, and then later resolve, ambiguity in their conceptual models. We leverage these findings to develop rTisane, a DSL for expressing conceptual models augmented with an interactive disambiguation process. In a controlled evaluation, we find that analysts reconsidered their assumptions, self-reported externalizing their assumptions accurately, and maintained analysis intent with rTisane. Additionally, rTisane enabled some analysts to author statistical models they were unable to specify manually. For others, rTisane resulted in models that better fit the data or enabled iterative improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64411195-7634-4e5a-8a2c-5d3d19e8839cCited by top-tier papers2
- PriorWeaver: Prior Elicitation via Iterative Dataset ConstructionYuwei Xiao, Shuai Ma, Antti Oulasvirta, Eunice JunCHI 2026 · 1 citation
- Causality and Semantic SeparationAnna Zhang, Qinglan Luo, London Bielicke, Eunice Jun et al.PLDI 2026
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- "I Don't Even Remember What I Read": How Design Influences Dissociation on Social MediaAmanda Baughan, Mingrui Ray Zhang, Raveena Rao, Kai Lukoff et al.CHI 2022 · 68 citations
- Paths Explored, Paths Omitted, Paths Obscured: Decision Points & Selective Reporting in End-to-End Data AnalysisYang Liu, Tim Althoff, Jeffrey HeerCHI 2020 · 40 citations
- Can You Hear My Heartbeat?: Hearing an Expressive Biosignal Elicits EmpathyR. Michael Winters, Bruce N. Walker, Grace LeslieCHI 2021 · 38 citations
- Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data RelationshipsEunice Jun, Audrey Seo, Jeffrey Heer, René JustCHI 2022 · 24 citations
Related papers
- Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling TasksVenkatesh Sivaraman, Anika Vaishampayan, Xiaotong Li, Brian R. Buck et al.CHI 2025 · 1 citation
- AutoDSL: Automated domain-specific language design for structural representation of procedures with constraintsYu-Zhe Shi, Haofei Hou, Zhangqian Bi, Fanxu Meng et al.ACL 2024
- How Do Analysts Understand and Verify AI-Assisted Data Analyses?Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang et al.CHI 2024 · 36 citations
- FlowNL: Asking the Flow Data in Natural LanguagesJieying Huang, Yang Xi, Junnan Hu, Jun TaoIEEE VIS 2022 · 13 citations
- EVM: Incorporating Model Checking into Exploratory Visual AnalysisAlex Kale, Ziyang Guo, Xiaoli Qiao, Jeffrey Heer et al.IEEE VIS 2023 · 16 citations
