FormaT5: Abstention and Examples for Conditional Table Formatting with Natural Language
Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le, Carina Negreanu, Elnaz Nouri, Mohammad Raza, Gust Verbruggen
摘要
Formatting is an important property in tables for visualization, presentation, and analysis. Spreadsheet software allows users to automatically format their tables by writing data-dependent conditional formatting (CF) rules. Writing such rules is often challenging for users as it requires understanding and implementing the underlying logic. We present FormaT5, a transformer-based model that can generate a CF rule given the target table and a natural language description of the desired formatting logic. We find that user descriptions for these tasks are often under-specified or ambiguous, making it harder for code generation systems to accurately learn the desired rule in a single step. To tackle this problem of under-specification and minimise argument errors, FormaT5 learns to predict placeholders though an abstention objective. These placeholders can then be filled by a second model or, when examples of rows that should be formatted are available, by a programming-by-example system. To evaluate FormaT5 on diverse and real scenarios, we create an extensive benchmark of 1053 CF tasks, containing real-world descriptions collected from four different sources. We release our benchmarks to encourage research in this area. Abstention and filling allow FormaT5 to outperform 8 different neural approaches on our benchmarks, both with and without examples. Our results illustrate the value of building domain-specific learning systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FLAME: A Small Language Model for Spreadsheet FormulasHarshit Joshi, Abishai Ebenezer, José Pablo Cambronero Sánchez, Sumit Gulwani 等AAAI 2024 · 被引用 21 次
- SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language ModelsYibin Chen, Yifu Yuan, Zeyu Zhang, Yan Zheng 等WWW 2025 · 被引用 13 次
- CoLA: Model Collaboration for Log-based Anomaly DetectionXuhang Zhu, Xiu Tang, Sai Wu, Jichen Li 等VLDB 2025 · 被引用 2 次
- PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language ModelsXianghong Xu, Xiao He, Tieying Zhang, Lei Zhang 等SIGMOD 2025 · 被引用 1 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-CorrectionMohammadreza Pourreza, Davood RafieiNeurIPS 2023 · 被引用 909 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi 等ICLR 2022 · 被引用 347 次
相关 Paper
- CORNET: Learning Table Formatting Rules By ExampleMukul Singh, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le 等VLDB 2023 · 被引用 11 次
- SpreadsheetCoder: Formula Prediction from Semi-structured ContextXinyun Chen, Petros Maniatis, Rishabh Singh, Charles Sutton 等ICML 2021 · 被引用 63 次
- ReFF: Reinforcing Format Faithfulness in Language Models Across Varied TasksJiashu Yao, Heyan Huang, Zeming Liu, Haoyu Wen 等AAAI 2025 · 被引用 1 次
- Numerical Formula Recognition from TablesQingping Yang, Yixuan Cao, Hongwei Li, Ping LuoKDD 2021 · 被引用 3 次
- PrepBench: How Far Are We from Natural-Language-Driven Data Preparation?Jingzhe Xu, Rui Wang, Jiannan Wang, Guoliang LiVLDB 2026 · 被引用 3 次
