Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts
Yunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang, Haotian Li, Huamin Qu, Di Weng, Yingcai Wu
Abstract
Data analysts frequently employ code completion tools in writing custom scripts to tackle complex tabular data wrangling tasks. However, existing tools do not sufficiently link the data contexts such as schemas and values with the code being edited. This not only leads to poor code suggestions, but also frequent interruptions in coding processes as users need additional code to locate and understand relevant data. We introduce Xavier, a tool designed to enhance data wrangling script authoring in computational notebooks. Xavier maintains users' awareness of data contexts while providing data-aware code suggestions. It automatically highlights the most relevant data based on the user's code, integrates both code and data contexts for more accurate suggestions, and instantly previews data transformation results for easy verification. To evaluate the effectiveness and usability of Xavier, we conducted a user study with 16 data analysts, showing its potential to streamline data wrangling scripts authoring.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3745ef75-1158-4e73-8c08-474648524c4fCited by top-tier papers3
- Cerebra: Aligning Implicit Knowledge in Interactive SQL AuthoringYunfan Zhou, Qiming Shi, Zhongsu Luo, Xiwen Cai et al.CHI 2026 · 2 citations
- Facilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeekYanwei Huang, Arpit NarechaniaCHI 2026 · 2 citations
- NoteFlow: Leveraging Charts as Sight Glasses for Consistent and Continuous Data Flow TracingYuan Tian, Dazhen Deng, Sen Yang, Huawei Zheng et al.CHI 2026 · 1 citation
Builds on29
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 260 citations
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 162 citations
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton et al.ICSE 2020 · 140 citations
Related papers
- Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsIan Drosos, Titus Barik, Philip J. Guo, Robert DeLine et al.CHI 2020 · 110 citations
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 25 citations
- Unravel: A Fluent Code Explorer for Data WranglingNischal Shrestha, Titus Barik, Chris ParninUIST 2021 · 19 citations
- Contextualized Data-Wrangling Code Generation in Computational NotebooksJunjie Huang, Daya Guo, Chenglong Wang, Jiazhen Gu et al.ASE 2024 · 5 citations
- Sneak Pique: Exploring Autocompletion as a Data Discovery Scaffold for Supporting Visual AnalysisVidya Setlur, Enamul Hoque, Dae Hyun Kim, Angel X. ChangUIST 2020 · 23 citations
