Ferry: Toward Better Understanding of Input/Output Space for Data Wrangling Scripts
Zhongsu Luo, Kai Xiong, Jiajun Zhu, Ran Chen, Xinhuan Shu, Di Weng, Yingcai Wu
摘要
Understanding the input and output of data wrangling scripts is crucial for various tasks like debugging code and onboarding new data. However, existing research on script understanding primarily focuses on revealing the process of data transformations, lacking the ability to analyze the potential scope, i.e., the space of script inputs and outputs. Meanwhile, constructing input/output space during script analysis is challenging, as the wrangling scripts could be semantically complex and diverse, and the association between different data objects is intricate. To facilitate data workers in understanding the input and output space of wrangling scripts, we summarize ten types of constraints to express table space and build a mapping between data transformations and these constraints to guide the construction of the input/output for individual transformations. Then, we propose a constraint generation model for integrating table constraints across multiple transformations. Based on the model, we develop Ferry, an interactive system that extracts and visualizes the data constraints describing the input and output space of data wrangling scripts, thereby enabling users to grasp the high-level semantics of complex scripts and locate the origins of faulty data transformations. Besides, Ferry provides example input and output data to assist users in interpreting the extracted constraints and checking and resolving the conflicts between these constraints and any uploaded dataset. Ferry's effectiveness and usability are evaluated through two usage scenarios and two case studies, including understanding, debugging, and checking both single and multiple scripts, with and without executable data. Furthermore, an illustrative application is presented to demonstrate Ferry's flexibility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling ScriptsYunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang 等CHI 2025 · 被引用 4 次
- ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMsYuan Tian, Chuhan Zhang, Xiaotong Wang, Sitong Pan 等UIST 2025 · 被引用 3 次
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User PromptsJiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou 等UIST 2025 · 被引用 1 次
它引用的顶会 Paper11
- Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsIan Drosos, Titus Barik, Philip J. Guo, Robert DeLine 等CHI 2020 · 被引用 110 次
- Data Formulator: AI-Powered Concept-Driven Visualization AuthoringChenglong Wang, John Thompson, Bongshin LeeIEEE VIS 2023 · 被引用 35 次
- Datamations: Animated Explanations of Data Analysis PipelinesXiaoying Pu, Sean Kross, Jake M. Hofman, Daniel G. GoldsteinCHI 2021 · 被引用 33 次
- Table Scraps: An Actionable Framework for Multi-Table Data Wrangling From An Artifact Study of Computational JournalismStephen Kasica, Charles Berret, Tamara MunznerIEEE VIS 2020 · 被引用 32 次
- Auto-Pipeline: Synthesize Data Pipelines By-Target Using Reinforcement Learning and SearchJunwen Yang, Yeye He, Surajit ChaudhuriVLDB 2021 · 被引用 32 次
相关 Paper
- Revealing the Semantics of Data Wrangling Scripts With ComanticsKai Xiong, Zhongsu Luo, Siwei Fu, Yongheng Wang 等IEEE VIS 2022 · 被引用 13 次
- Unravel: A Fluent Code Explorer for Data WranglingNischal Shrestha, Titus Barik, Chris ParninUIST 2021 · 被引用 19 次
- Dango: A Mixed-Initiative Data Wrangling System using Large Language ModelWei-Hao Chen, Weixi Tong, Amanda Case, Tianyi ZhangCHI 2025 · 被引用 19 次
- VizLinter: A Linter and Fixer Framework for Data VisualizationQing Chen, Fuling Sun, Xinyue Xu, Zui Chen 等IEEE VIS 2021 · 被引用 60 次
- Diagramming Program Values by Spatial RefinementSiddhartha Prasad, Michael Tu, Karan Kashyap, Tim Nelson 等PLDI 2026
