Database reasoning over text
James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, Alon Y. Halevy
Abstract
Neural models have shown impressive performance gains in answering queries from natural language text. However, existing works are unable to support database queries, such as "List/Count all female athletes who were born in 20th century", which require reasoning over sets of relevant facts with operations such as join, filtering and aggregation. We show that while state-of-the-art transformer models perform very well for small databases, they exhibit limitations in processing noisy data, numerical operations, and queries that aggregate facts. We propose a modular architecture to answer these database-style queries over multiple spans from text and aggregating these at scale. We evaluate the architecture using WIKINLDB, 1 a novel dataset for exploring such queries. Our architecture scales to databases containing thousands of facts whereas contemporary models are limited by how many facts can be encoded. In direct comparison on small databases, our approach increases overall answer accuracy from 85% to 90%. On larger databases, our approach retains its accuracy whereas transformer baselines could not encode the context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88ce164e-e725-4f8e-b760-01e78c2fb30fCited by top-tier papers7
- Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsOri Yoran, Alon Talmor, Jonathan BerantACL 2022 · 57 citations
- Binding Language Models in Symbolic LanguagesZhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li et al.ICLR 2023 · 38 citations
- Relational World Knowledge Representation in Contextual Language Models: A ReviewTara Safavi, Danai KoutraEMNLP 2021 · 31 citations
- ELEET: Efficient Learned Query Execution over Text and TablesMatthias Urban, Carsten BinnigVLDB 2024 · 15 citations
- A Data Source for Reasoning Embodied AgentsJack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve, Yuxuan Sun et al.AAAI 2023 · 10 citations
Builds on4
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Answering Complex Open-Domain Questions with Multi-Hop Dense RetrievalWenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du et al.ICLR 2021 · 232 citations
- Neural Module Networks for Reasoning over TextNitish Gupta, Kevin Lin, Dan Roth, Sameer Singh et al.ICLR 2020 · 134 citations
- From Natural Language Processing to Neural DatabasesJames Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri et al.VLDB 2021 · 62 citations
Related papers
- Weaver: Interweaving SQL and LLM for Table ReasoningRohit Khoja, Devanshu Gupta, Yanjie Fu, Dan Roth et al.EMNLP 2025 · 1 citation
- SPARQLing Database Queries from Intermediate Question DecompositionsIrina Saparina, Anton OsokinEMNLP 2021 · 11 citations
- KaggleDBQA: Realistic Evaluation of Text-to-SQL ParsersChia-Hsuan Lee, Oleksandr Polozov, Matthew RichardsonACL 2021
- RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersBailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov et al.ACL 2020 · 39 citations
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in LLMsSoyeon Kim, Jindong Wang, Xing Xie, Steven Euijong WhangICLR 2026
