MATE: Multi-view Attention for Table Transformer Efficiency
Julian Martin Eisenschlos, Maharshi Gor, Thomas Müller, William W. Cohen
摘要
This work presents a sparse-attention Transformer architecture for modeling documents that contain large tables. Tables are ubiquitous on the web, and are rich in information. However, more than 20% of relational tables on the web have 20 or more rows (Cafarella et al., 2008) , and these large tables present a challenge for current Transformer models, which are typically limited to 512 tokens. Here we propose MATE, a novel Transformer architecture designed to model the structure of web tables. MATE uses sparse attention in a way that allows heads to efficiently attend to either rows or columns in a table. This architecture scales linearly with respect to speed and memory, and can handle documents containing more than 8000 tokens with current accelerators. MATE also has a more appropriate inductive bias for tabular data, and sets a new state-of-the-art for three table reasoning datasets. For HY-BRIDQA (Chen et al., 2020b), a dataset that involves large documents containing tables, we improve the best prior result by 19 points. * Work done at Google Research. 1 The term "semi-structured text" refers to text that has structure that does not reflect a known data schema. Typically semi-structured text is organized as an HTML tree or variable length lists and tables.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong 等EMNLP 2022 · 被引用 222 次
- TableFormer: Robust Transformer Modeling for Table-Text EncodingJingfeng Yang, Aditya Gupta, Shyam Upadhyay, Luheng He 等ACL 2022 · 被引用 145 次
- UL2: Unifying Language Learning ParadigmsYi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia 等ICLR 2023 · 被引用 97 次
- HyTrel: Hypergraph-enhanced Tabular Data Representation LearningPei Chen, Soumajyoti Sarkar, Leonard Lausen, Balasubramaniam Srinivasan 等NeurIPS 2023 · 被引用 66 次
- UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data ScienceYazheng Yang, Yuqi Wang, Guang Liu, Ledell Wu 等ICLR 2024 · 被引用 35 次
它引用的顶会 Paper10
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek 等EMNLP 2020 · 被引用 268 次
相关 Paper
- Multi-Row, Multi-Span Distant Supervision For Table+Text Question AnsweringVishwajeet Kumar, Yash Gupta, Saneem A. Chemmengath, Jaydeep Sen 等ACL 2023 · 被引用 3 次
- TempTabQA: Temporal Question Answering for Semi-Structured TablesVivek Gupta, Pranshu Kandoi, Mahek Bhavesh Vora, Shuo Zhang 等EMNLP 2023 · 被引用 4 次
- Towards Cross-Table Masked Pretraining for Web Data MiningChao Ye, Guoshan Lu, Haobo Wang, Liyao Li 等WWW 2024 · 被引用 23 次
- Table as a Modality for Large Language ModelsLiyao Li, Chao Ye, Wentao Ye, Yifei Sun 等NeurIPS 2025 · 被引用 5 次
- GetPt: Graph-enhanced General Table Pre-training with Alternate Attention NetworkRan Jia, Haoming Guo, Xiaoyuan Jin, Chao Yan 等KDD 2023 · 被引用 3 次
