Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning Skills
Ori Yoran, Alon Talmor, Jonathan Berant
摘要
Models pre-trained with a language modeling objective possess ample world knowledge and language skills, but are known to struggle in tasks that require reasoning. In this work, we propose to leverage semi-structured tables, and automatically generate at scale questionparagraph pairs, where answering the question requires reasoning over multiple facts in the paragraph. We add a pre-training step over this synthetic data, which includes examples that require 16 different reasoning skills such as number comparison, conjunction, and fact composition. To improve data efficiency, we propose sampling strategies that focus training on reasoning skills the model is currently lacking. We evaluate our approach on three reading comprehension datasets that are focused on reasoning, and show that our model, PReasM, substantially outperforms T5, a popular pre-trained encoder-decoder model. Moreover, sampling examples based on current model errors leads to faster training and higher overall performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong 等EMNLP 2022 · 被引用 222 次
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa 等AAAI 2023 · 被引用 178 次
- HyTrel: Hypergraph-enhanced Tabular Data Representation LearningPei Chen, Soumajyoti Sarkar, Leonard Lausen, Balasubramaniam Srinivasan 等NeurIPS 2023 · 被引用 66 次
- Making Language Models Better Reasoners with Step-Aware VerifierYifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu 等ACL 2023 · 被引用 52 次
- On Grounded Planning for Embodied Tasks with Language ModelsBill Yuchen Lin, Chengsong Huang, Qian Liu, Wenda Gu 等AAAI 2023 · 被引用 52 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
相关 Paper
- ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning ExamplesYilun Zhao, Linyong Nan, Zhenting Qi, Rui Zhang 等EMNLP 2022 · 被引用 15 次
- Probing How Scalable Table Data Enhances General Long-Context ReasoningHuaibing Xie, Guoliang Zhao, Yang Liu, Shihan Dou 等ICML 2026
- CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular TablesZhen Yang, Wei Du, Jie Wang, Wenze Zhou 等ACL 2026
- Pre-training Language Models for Comparative ReasoningMengxia Yu, Zhihan Zhang, Wenhao Yu, Meng JiangEMNLP 2023 · 被引用 1 次
- Leap-Of-Thought: Teaching Pre-Trained Models to Systematically Reason Over Implicit KnowledgeAlon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg 等NeurIPS 2020 · 被引用 119 次
