TabulaX: Leveraging Large Language Models for Multi-Class Table Transformations
Arash Dargahi Nobari, Davood Rafiei
Abstract
The integration of tabular data from diverse sources is often hindered by inconsistencies in formatting and representation, posing significant challenges for data analysts and personal digital assistants. Existing methods for automating tabular data transformations are limited in scope, often focusing on specific types of transformations or lacking interpretability. In this paper, we introduce TabulaX, a novel framework that leverages Large Language Models (LLMs) for multi-class column-level tabular transformations. TabulaX first classifies input columns into four transformation types—string-based, numerical, algorithmic, and general—and then applies tailored methods to generate human-interpretable transformation functions, such as numeric formulas or programming code. This approach enhances transparency and allows users to understand and modify the mappings. Through extensive experiments on real-world datasets from various domains, we demonstrate that TabulaX outperforms existing state-of-the-art approaches in terms of accuracy, supports a broader class of transformations, and generates interpretable transformations that can be efficiently applied.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b970d8e-33ad-41f2-bc60-330a64414023Cited by top-tier papers3
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex DocumentsShu Wang, Yingli Zhou, Yixiang FangVLDB 2026 · 16 citations
- SQL-Exchange: Transforming SQL Queries Across DomainsMohammadreza Daviran, Brian Lin, Davood RafieiVLDB 2026 · 2 citations
- EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language QueriesYuhui Wang, Jinqi Liu, Chengliang Chai, Hangyu Zhao et al.VLDB 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Visualization Recommendation with Prompt-based Reprogramming of Large Language ModelsXinhang Li, Jingbo Zhou, Wei Chen, Derong Xu et al.ACL 2024 · 4 citations
- Reasoning and Retrieval for Complex Semi-structured Tables via Reinforced Relational Data TransformationHaoyu Dong, Yue Hu, Yanan CaoSIGIR 2025 · 2 citations
- AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent FrameworkMeihao Fan, Ju Fan, Nan Tang, Lei Cao et al.VLDB 2025 · 10 citations
- DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language ModelsArash Dargahi Nobari, Davood RafieiSIGMOD 2024 · 11 citations
- Empowering Tabular Data Preparation with Language Models: Why and How?Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang et al.ACL 2026 · 4 citations
