Xavier: Toward Better Coding Assistance in Authoring Tabular Data Wrangling Scripts
Yunfan Zhou, Xiwen Cai, Qiming Shi, Yanwei Huang, Haotian Li, Huamin Qu, Di Weng, Yingcai Wu
摘要
Data analysts frequently employ code completion tools in writing custom scripts to tackle complex tabular data wrangling tasks. However, existing tools do not sufficiently link the data contexts such as schemas and values with the code being edited. This not only leads to poor code suggestions, but also frequent interruptions in coding processes as users need additional code to locate and understand relevant data. We introduce Xavier, a tool designed to enhance data wrangling script authoring in computational notebooks. Xavier maintains users' awareness of data contexts while providing data-aware code suggestions. It automatically highlights the most relevant data based on the user's code, integrates both code and data contexts for more accurate suggestions, and instantly previews data transformation results for easy verification. To evaluate the effectiveness and usability of Xavier, we conducted a user study with 16 data analysts, showing its potential to streamline data wrangling scripts authoring.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Cerebra: Aligning Implicit Knowledge in Interactive SQL AuthoringYunfan Zhou, Qiming Shi, Zhongsu Luo, Xiwen Cai 等CHI 2026 · 被引用 2 次
- Facilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeekYanwei Huang, Arpit NarechaniaCHI 2026 · 被引用 2 次
- NoteFlow: Leveraging Charts as Sight Glasses for Consistent and Continuous Data Flow TracingYuan Tian, Dazhen Deng, Sen Yang, Huawei Zheng 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper29
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 被引用 408 次
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 被引用 162 次
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton 等ICSE 2020 · 被引用 140 次
相关 Paper
- Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsIan Drosos, Titus Barik, Philip J. Guo, Robert DeLine 等CHI 2020 · 被引用 110 次
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 被引用 25 次
- Unravel: A Fluent Code Explorer for Data WranglingNischal Shrestha, Titus Barik, Chris ParninUIST 2021 · 被引用 19 次
- Contextualized Data-Wrangling Code Generation in Computational NotebooksJunjie Huang, Daya Guo, Chenglong Wang, Jiazhen Gu 等ASE 2024 · 被引用 5 次
- Sneak Pique: Exploring Autocompletion as a Data Discovery Scaffold for Supporting Visual AnalysisVidya Setlur, Enamul Hoque, Dae Hyun Kim, Angel X. ChangUIST 2020 · 被引用 23 次
