PixT3: Pixel-based Table-To-Text Generation
Iñigo Alonso, Eneko Agirre, Mirella Lapata
摘要
Table-to-text generation involves generating appropriate textual descriptions given structured tabular data. It has attracted increasing attention in recent years thanks to the popularity of neural network models and the availability of large-scale datasets. A common feature across existing methods is their treatment of the input as a string, i.e., by employing linearization techniques that do not always preserve information in the table, are verbose, and lack space efficiency. We propose to rethink data-to-text generation as a visual recognition task, removing the need for rendering the input in a string format. We present PixT3, a multimodal tableto-text model that overcomes the challenges of linearization and input size limitations encountered by existing models. PixT3 is trained with a new self-supervised learning objective to reinforce table structure awareness and is applicable to open-ended and controlled generation settings. Experiments on the ToTTo (Parikh et al., 2020a) and Logic2Text (Chen et al., 2020c) benchmarks show that PixT3 is competitive and, in some settings, superior to generators that operate solely on text. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TABLET: A Large-Scale Dataset for Robust Visual Table UnderstandingIñigo Alonso, Imanol Miranda, Eneko Agirre, Mirella LapataICLR 2026 · 被引用 5 次
- GRIT: Guided Relational Integration for Efficient Multi-Table UnderstandingYujin Kang, Park Seong Woo, Yoon-Sik ChoEMNLP 2025 · 被引用 2 次
- QuASAR: A Question-Driven Structure-Aware Approach for Table-to-Text GenerationWeijie Liu, Yibin Zheng, Fang KongACL 2025
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
相关 Paper
- TableVLM: Multi-modal Pre-training for Table Structure RecognitionLeiyuan Chen, Chengsong Huang, Xiaoqing Zheng, Jinshu Lin 等ACL 2023 · 被引用 8 次
- VisToT: Vision-Augmented Table-to-Text GenerationPrajwal Gatti, Anand Mishra, Manish Gupta, Mithun Das GuptaEMNLP 2022 · 被引用 4 次
- Bridging the Semantic Gap Between Text and Table: A Case Study on NL2SQLLin Long, Xijun Gu, Xinjie Sun, Wentao Ye 等ICLR 2025
- PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image SegmentationViktor Zaytsev, Olena Vynokurova, Pavlo Tytarchuk, Dmytro Kozii 等CVPR 2026
- Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source LearningAlexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 等ACL 2023 · 被引用 2 次
