PixT3: Pixel-based Table-To-Text Generation
Iñigo Alonso, Eneko Agirre, Mirella Lapata
Abstract
Table-to-text generation involves generating appropriate textual descriptions given structured tabular data. It has attracted increasing attention in recent years thanks to the popularity of neural network models and the availability of large-scale datasets. A common feature across existing methods is their treatment of the input as a string, i.e., by employing linearization techniques that do not always preserve information in the table, are verbose, and lack space efficiency. We propose to rethink data-to-text generation as a visual recognition task, removing the need for rendering the input in a string format. We present PixT3, a multimodal tableto-text model that overcomes the challenges of linearization and input size limitations encountered by existing models. PixT3 is trained with a new self-supervised learning objective to reinforce table structure awareness and is applicable to open-ended and controlled generation settings. Experiments on the ToTTo (Parikh et al., 2020a) and Logic2Text (Chen et al., 2020c) benchmarks show that PixT3 is competitive and, in some settings, superior to generators that operate solely on text. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa915254-f71a-4e46-9a49-788384283fd0Cited by top-tier papers3
- TABLET: A Large-Scale Dataset for Robust Visual Table UnderstandingIñigo Alonso, Imanol Miranda, Eneko Agirre, Mirella LapataICLR 2026 · 5 citations
- GRIT: Guided Relational Integration for Efficient Multi-Table UnderstandingYujin Kang, Park Seong Woo, Yoon-Sik ChoEMNLP 2025 · 2 citations
- QuASAR: A Question-Driven Structure-Aware Approach for Table-to-Text GenerationWeijie Liu, Yibin Zheng, Fang KongACL 2025
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
Related papers
- TableVLM: Multi-modal Pre-training for Table Structure RecognitionLeiyuan Chen, Chengsong Huang, Xiaoqing Zheng, Jinshu Lin et al.ACL 2023 · 8 citations
- VisToT: Vision-Augmented Table-to-Text GenerationPrajwal Gatti, Anand Mishra, Manish Gupta, Mithun Das GuptaEMNLP 2022 · 4 citations
- Bridging the Semantic Gap Between Text and Table: A Case Study on NL2SQLLin Long, Xijun Gu, Xinjie Sun, Wentao Ye et al.ICLR 2025
- PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image SegmentationViktor Zaytsev, Olena Vynokurova, Pavlo Tytarchuk, Dmytro Kozii et al.CVPR 2026
- Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source LearningAlexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma et al.ACL 2023 · 2 citations
