CORNET: Learning Table Formatting Rules By Example
Mukul Singh, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le, Carina Negreanu, Mohammad Raza, Gust Verbruggen
Abstract
Spreadsheets are widely used for table manipulation and presentation. Stylistic formatting of these tables is an important property for presentation and analysis. As a result, popular spreadsheet software, such as Excel, supports automatically formatting tables based on rules. Unfortunately, writing such formatting rules can be challenging for users as it requires knowledge of the underlying rule language and data logic. We present Cornet, a system that tackles the novel problem of automatically learning such formatting rules from user-provided formatted cells. Cornet takes inspiration from advances in inductive programming and combines symbolic rule enumeration with a neural ranker to learn conditional formatting rules. To motivate and evaluate our approach, we extracted tables with over 450K unique formatting rules from a corpus of over 1.8M real worksheets. Since we are the first to introduce the task of automatically learning conditional formatting rules, we compare Cornet to a wide range of symbolic and neural baselines adapted from related domains. Our results show that Cornet accurately learns rules across varying setups. Additionally, we show that in some cases Cornet can find rules that are shorter than those written by users and can also discover rules in spreadsheets that users have manually formatted. Furthermore, we present two case studies investigating the generality of our approach by extending Cornet to related data tasks (e.g., filtering) and generalizing to conditional formatting over multiple columns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Synchromesh: Reliable Code Generation from Pre-trained Language ModelsGabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari et al.ICLR 2022 · 200 citations
- Wrex: A Unified Programming-by-Example Interaction for Synthesizing Readable Code for Data ScientistsIan Drosos, Titus Barik, Philip J. Guo, Robert DeLine et al.CHI 2020 · 110 citations
- TUTA: Tree-based Transformers for Generally Structured Table Pre-trainingZhiruo Wang, Haoyu Dong, Ran Jia, Jia Li et al.KDD 2021 · 88 citations
- SpreadsheetCoder: Formula Prediction from Semi-structured ContextXinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton et al.ICML 2021 · 63 citations
Related papers
- FormaT5: Abstention and Examples for Conditional Table Formatting with Natural LanguageMukul Singh, José Cambronero, Sumit Gulwani, Vu Le et al.VLDB 2024 · 13 citations
- Auto-Formula: Recommend Formulas in Spreadsheets using Contrastive Learning for Table RepresentationsSibei Chen, Yeye He, Weiwei Cui, Ju Fan et al.SIGMOD 2024 · 4 citations
- Seq2Parse: neurosymbolic parse error repairGeorgios Sakkas, Madeline Endres, Philip J. Guo, Westley Weimer et al.OOPSLA 2022 · 6 citations
- Numerical Formula Recognition from TablesQingping Yang, Yixuan Cao, Hongwei Li, Ping LuoKDD 2021 · 3 citations
- "It's Freedom to Put Things Where My Mind Wants": Understanding and Improving the User Experience of Structuring Data in SpreadsheetsGeorge Chalhoub, Advait SarkarCHI 2022 · 21 citations
