Efficient and Scalable Search for Statistics
Antoine Gauquier, Simon Ebel, Helena Galhardas, Théo Galizzi, Ioana Manolescu, Aurélien Peden, Pierre Senellart
摘要
Informed public debate needs high-quality data. In this context, high-quality statistical data sources are a valuable category of reference information based on which a claim can be checked. To facilitate the work of journalists or other factcheckers, users' questions about a specific claim should be automatically answered based on statistical tables. This task is complicated by the large number, size, and variety of statistical datasets. We introduce the statistical table discovery problem (STD, in short), which aims, given a natural language question and a set of statistical datasets (multidimensional tables), to find the tables most relevant for the question. We then describe STAR, an algorithm for solving the STD problem. Unlike existing table discovery (TD) solutions aimed at relational tables, STAR is devised specifically for multidimensional ones. Further, STAR treats the space and time dimensions of statistical datasets separately. We experimentally show that these features, together, make STAR outperform state-of-the-art TD systems adapted to the STD problem, in terms of scalability, search quality, preprocessing and question answering time.
short), possibly fine-tuned on these serializations [8]. Such methods, by design, reflect (or learn) the serialization order of cells in a row, yet this order is meaningless in statistical tables. As a consequence, result quality may be negatively impacted.
A second core insight in StatCheck [6], [7], is that the fact values, which make up a majority of cells in statistical tables, answer questions but are not part of the questions themselves, for TD nor for TQA. Only the table cells other than fact values need to be pre-processed (indexed), so that at query time the index can lead to the fact values that answer the question. As we will show, this keeps StatCheck indexes of reasonable size even on large statistical corpora, on which indexing is unfeasible (taking excessive time or space) or much slower for competitor systems [2]-[4]. Finally, the novel insight of [7] is: statistical data dimensions, time and space, are omnipresent in statistical data, and frequent also in user queries. Therefore, the Spaceand Time aware STatistic Retrieval (STAR, in short) method solves the TD problem by extracting space and time dimension values from the statistics, and from user queries, and separately processing the time, space, and remaining part of the search query. As shown in [7], this has significantly improved the quality of StatCheck results.
In this work, we make the following contributions.
• We formalize the new Statistical Table Discovery (STD, in short) problem as a variant of TD, adapted to the particularities of multidimensional statistical tables.
• We detail the STAR method for the STD problem, only briefly outlined in the demonstration [7]. While the demonstration provided limited evidence for some core choices in STAR, this paper details it, and thoroughly compares it to state-of-the-art TD systems on the STD task, on a new benchmark (see below).
• We improve over [7] by replacing the popular SBert [9] phrase embedding model used in [7] with the recent PEARL [10] model better adapted to short phrases; this improves result quality by up to 14%. We introduce a relaxation mechanism that allows STAR to return tables semantically relevant even when no exact match exists for the specified time or location. Additionally, we extend STAR to handle natural language questions, which are more suitable than keywords for complex queries and pave the way for combining STAR with TQA systems for full Statistical Question Answering.
• We generalize STAR's score to reflect the distance between a user question and a candidate statistical table, along three dimensions: space, time, and the measure itself, enabling it to identify more answers that are relevant for a given question.
• We present a novel, comprehensive performance comparison on the STD task, between STAR, the recent systems [2]-[4], and the classic BM25 retrieval metric.
Our experiments demonstrate that: (i) when user questions use the exact terminology of the datasets, the BM25 method outperforms all the others in terms of result quality, while also providing fast indexing and query answering; (ii) when questions are reformulated to use
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Distilling Knowledge from Reader to Retriever for Question AnsweringGautier Izacard, Edouard GraveICLR 2021 · 被引用 317 次
- Table Search Using a Deep Contextualized Language ModelZhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu 等SIGIR 2020 · 被引用 48 次
- Web Table Retrieval using Multimodal Deep LearningRoee Shraga, Haggai Roitman, Guy Feigenblat, Mustafa CanimSIGIR 2020 · 被引用 45 次
- Solo: Data Discovery Using Natural Language Questions Via A Self-Supervised ApproachQiming Wang, Raul Castro FernandezSIGMOD 2024 · 被引用 20 次
- TaPas: Weakly Supervised Table Parsing via Pre-trainingJonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno 等ACL 2020 · 被引用 19 次
相关 Paper
- Extracting Contextualized Quantity Facts from Web TablesVinh Thinh Ho, Koninika Pal, Simon Razniewski, Klaus Berberich 等WWW 2021 · 被引用 13 次
- HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language GenerationZhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia 等ACL 2022
- Gen-T: Table Reclamation in Data LakesGrace Fan, Roee Shraga, Renée J. MillerICDE 2024 · 被引用 5 次
- Revisiting Single-Table Retrieval: An Open Problem Under 360° Stress TestsChenyu Yang, Ziyu Jiang, Junhao Li, Yuyu Luo 等ICDE 2026
- Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question AnsweringFeng Luo, Hai Lan, Hui Luo, Zhifeng Bao 等ICDE 2026 · 被引用 1 次
