Data Formulator 2: Iterative Creation of Data Visualizations, with AI Transforming Data Along the Way
Chenglong Wang, Bongshin Lee, Steven Mark Drucker, Dan Marshall, Jianfeng Gao
Abstract
Data analysts often need to iterate between data transformations and chart designs to create rich visualizations for exploratory data analysis. Although many AI-powered systems have been introduced to reduce the effort of visualization authoring, existing systems are not well suited for iterative authoring. They typically require analysts to provide, in a single turn, a text-only prompt that fully describe a complex visualization. We introduce Data Formulator 2 (Df2 for short), an AI-powered visualization system designed to overcome this limitation. Df2 blends graphical user interfaces and natural language inputs to enable users to convey their intent more effectively, while delegating data transformation to AI. Furthermore, to support efficient iteration, Df2 lets users navigate their iteration history and reuse previous designs, eliminating the need to start from scratch each time. A user study with eight participants demonstrated that Df2 allowed participants to develop their own iteration styles to complete challenging data exploration sessions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d10b6cd-e141-4a4c-a6a9-21974edb86f5Cited by top-tier papers2
- DataWink: Reusing and Adapting SVG-Based Visualization Examples with Large Multimodal ModelsLiwenhan Xie, Yanna Lin, Can Liu, Huamin Qu et al.IEEE VIS 2025 · 3 citations
- CoFEH: LLM-driven Feature Engineering Empowered by Collaborative Bayesian Hyperparameter OptimizationBeicheng Xu, Keyao Ding, Wei Liu, Yupeng Lu et al.KDD 2026
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
Related papers
- Data Formulator: AI-Powered Concept-Driven Visualization AuthoringChenglong Wang, John Thompson, Bongshin LeeIEEE VIS 2023 · 35 citations
- Towards Natural Language-Based Visualization AuthoringYun Wang, Zhitao Hou, Leixian Shen, Tongshuang Wu et al.IEEE VIS 2022 · 74 citations
- Telling Stories from Computational Notebooks: AI-Assisted Presentation Slides Creation for Presenting Data Science WorkChengbo Zheng, Dakuo Wang, April Yi Wang, Xiaojuan MaCHI 2022 · 53 citations
- InfoAlign: A Human-AI Co-Creation System for Storytelling with InfographicsJielin Feng, Xinwu Ye, Qianhui Li, Verena Ingrid Prantl et al.CHI 2026 · 1 citation
- Sneak Pique: Exploring Autocompletion as a Data Discovery Scaffold for Supporting Visual AnalysisVidya Setlur, Enamul Hoque, Dae Hyun Kim, Angel X. ChangUIST 2020 · 23 citations
