Lune

ICDE2026顶会

Efficient and Scalable Search for Statistics

Antoine Gauquier, Simon Ebel, Helena Galhardas, Théo Galizzi, Ioana Manolescu, Aurélien Peden, Pierre Senellart

2026年份

摘要

Informed public debate needs high-quality data. In this context, high-quality statistical data sources are a valuable category of reference information based on which a claim can be checked. To facilitate the work of journalists or other factcheckers, users' questions about a specific claim should be automatically answered based on statistical tables. This task is complicated by the large number, size, and variety of statistical datasets. We introduce the statistical table discovery problem (STD, in short), which aims, given a natural language question and a set of statistical datasets (multidimensional tables), to find the tables most relevant for the question. We then describe STAR, an algorithm for solving the STD problem. Unlike existing table discovery (TD) solutions aimed at relational tables, STAR is devised specifically for multidimensional ones. Further, STAR treats the space and time dimensions of statistical datasets separately. We experimentally show that these features, together, make STAR outperform state-of-the-art TD systems adapted to the STD problem, in terms of scalability, search quality, preprocessing and question answering time.

short), possibly fine-tuned on these serializations [8]. Such methods, by design, reflect (or learn) the serialization order of cells in a row, yet this order is meaningless in statistical tables. As a consequence, result quality may be negatively impacted.

A second core insight in StatCheck [6], [7], is that the fact values, which make up a majority of cells in statistical tables, answer questions but are not part of the questions themselves, for TD nor for TQA. Only the table cells other than fact values need to be pre-processed (indexed), so that at query time the index can lead to the fact values that answer the question. As we will show, this keeps StatCheck indexes of reasonable size even on large statistical corpora, on which indexing is unfeasible (taking excessive time or space) or much slower for competitor systems [2]-[4]. Finally, the novel insight of [7] is: statistical data dimensions, time and space, are omnipresent in statistical data, and frequent also in user queries. Therefore, the Spaceand Time aware STatistic Retrieval (STAR, in short) method solves the TD problem by extracting space and time dimension values from the statistics, and from user queries, and separately processing the time, space, and remaining part of the search query. As shown in [7], this has significantly improved the quality of StatCheck results.

In this work, we make the following contributions.

• We formalize the new Statistical Table Discovery (STD, in short) problem as a variant of TD, adapted to the particularities of multidimensional statistical tables.

• We detail the STAR method for the STD problem, only briefly outlined in the demonstration [7]. While the demonstration provided limited evidence for some core choices in STAR, this paper details it, and thoroughly compares it to state-of-the-art TD systems on the STD task, on a new benchmark (see below).

• We improve over [7] by replacing the popular SBert [9] phrase embedding model used in [7] with the recent PEARL [10] model better adapted to short phrases; this improves result quality by up to 14%. We introduce a relaxation mechanism that allows STAR to return tables semantically relevant even when no exact match exists for the specified time or location. Additionally, we extend STAR to handle natural language questions, which are more suitable than keywords for complex queries and pave the way for combining STAR with TQA systems for full Statistical Question Answering.

• We generalize STAR's score to reflect the distance between a user question and a candidate statistical table, along three dimensions: space, time, and the measure itself, enabling it to identify more answers that are relevant for a given question.

• We present a novel, comprehensive performance comparison on the STD task, between STAR, the recent systems [2]-[4], and the classic BM25 retrieval metric.

Our experiments demonstrate that: (i) when user questions use the exact terminology of the datasets, the BM25 method outperforms all the others in terms of result quality, while also providing fast indexing and query answering; (ii) when questions are reformulated to use

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4565d5cc-cb44-4cd2-9035-361e8ab8f37b

它引用的顶会 Paper6

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖