rTisane: Externalizing conceptual models for data analysis prompts reconsideration of domain assumptions and facilitates statistical modeling
Eunice Jun, Edward Misback, Jeffrey Heer, René Just
摘要
Statistical models should accurately reflect analysts’ domain knowledge about variables and their relationships. While recent tools let analysts express these assumptions and use them to produce a resulting statistical model, it remains unclear what analysts want to express and how externalization impacts statistical model quality. This paper addresses these gaps. We first conduct an exploratory study of analysts using a domain-specific language (DSL) to express conceptual models. We observe a preference for detailing how variables relate and a desire to allow, and then later resolve, ambiguity in their conceptual models. We leverage these findings to develop rTisane, a DSL for expressing conceptual models augmented with an interactive disambiguation process. In a controlled evaluation, we find that analysts reconsidered their assumptions, self-reported externalizing their assumptions accurately, and maintained analysis intent with rTisane. Additionally, rTisane enabled some analysts to author statistical models they were unable to specify manually. For others, rTisane resulted in models that better fit the data or enabled iterative improvement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PriorWeaver: Prior Elicitation via Iterative Dataset ConstructionYuwei Xiao, Shuai Ma, Antti Oulasvirta, Eunice JunCHI 2026 · 被引用 1 次
- Causality and Semantic SeparationAnna Zhang, Qinglan Luo, London Bielicke, Eunice Jun 等PLDI 2026
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- "I Don't Even Remember What I Read": How Design Influences Dissociation on Social MediaAmanda Baughan, Mingrui Ray Zhang, Raveena Rao, Kai Lukoff 等CHI 2022 · 被引用 68 次
- Paths Explored, Paths Omitted, Paths Obscured: Decision Points & Selective Reporting in End-to-End Data AnalysisYang Liu, Tim Althoff, Jeffrey HeerCHI 2020 · 被引用 40 次
- Can You Hear My Heartbeat?: Hearing an Expressive Biosignal Elicits EmpathyR. Michael Winters, Bruce N. Walker, Grace LeslieCHI 2021 · 被引用 38 次
- Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data RelationshipsEunice Jun, Audrey Seo, Jeffrey Heer, René JustCHI 2022 · 被引用 24 次
相关 Paper
- Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling TasksVenkatesh Sivaraman, Anika Vaishampayan, Xiaotong Li, Brian R. Buck 等CHI 2025 · 被引用 1 次
- AutoDSL: Automated domain-specific language design for structural representation of procedures with constraintsYu-Zhe Shi, Haofei Hou, Zhangqian Bi, Fanxu Meng 等ACL 2024
- How Do Analysts Understand and Verify AI-Assisted Data Analyses?Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang 等CHI 2024 · 被引用 36 次
- FlowNL: Asking the Flow Data in Natural LanguagesJieying Huang, Yang Xi, Junnan Hu, Jun TaoIEEE VIS 2022 · 被引用 13 次
- EVM: Incorporating Model Checking into Exploratory Visual AnalysisAlex Kale, Ziyang Guo, Xiaoli Qiao, Jeffrey Heer 等IEEE VIS 2023 · 被引用 16 次
