VisToT: Vision-Augmented Table-to-Text Generation
Prajwal Gatti, Anand Mishra, Manish Gupta, Mithun Das Gupta
Abstract
Table-to-text generation has been widely studied in the Natural Language Processing community in the recent years. We give a new perspective to this problem by incorporating signals from both tables as well as associated images to generate relevant text. While tables contain a structured list of facts, images are a rich source of unstructured visual information. For example, in the tourism domain, images can be used to infer knowledge such as the type of landmark (e.g., church), its architecture (e.g., Ancient Roman), and composition (e.g., white marble). Therefore, in this paper, we introduce the novel task of Vision-augmented Table-To-Text Generation (VISTOT), defined as follows: given a table and an associated image, produce a descriptive sentence conditioned on the multimodal input. For the task, we present a novel multimodal table-to-text dataset, WIKILAND-MARKS, covering 73,084 unique world landmarks. Further, we also present a competitive architecture, namely, VT3 that generates accurate sentences conditioned on the image and table pairs. Through extensive analyses and experiments, we show that visual cues from images are helpful in (i) inferring missing information from incomplete or sparse tables, and (ii) strengthening the importance of useful information from noisy tables for natural language generation. We make the code and data publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a8518fb-413a-423b-b29b-742197c07d2cBuilds on8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- End-to-End Transformer Based Model for Image CaptioningYiyu Wang, Jungang Xu, Yingfei SunAAAI 2022 · 178 citations
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen et al.ACL 2020 · 116 citations
Related papers
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
- Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal VisionXiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah SnavelyICCV 2021 · 26 citations
- Beyond Text-Only: Towards Multimodal Table Retrieval in Open-WorldDa Li, Keping Bi, Jiafeng Guo, Wei Yuan et al.ICLR 2026
- PixT3: Pixel-based Table-To-Text GenerationIñigo Alonso, Eneko Agirre, Mirella LapataACL 2024 · 2 citations
- Multimodal Table UnderstandingMingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She et al.ACL 2024
