Toward a Unified Framework for Unsupervised Complex Tabular Reasoning
Zhenyu Li, Xiuxing Li, Zhichao Duan, Bowen Dong, Ning Liu, Jianyong Wang
摘要
Structured tabular data exist across nearly all fields. Reasoning task over these data aims to answer questions or determine the truthiness of hypothesis sentences by understanding the semantic meaning of a table. While previous works have devoted significant efforts to the tabular reasoning task, they always assume there are sufficient labeled data. However, constructing reasoning samples over tables (and related text) is labor-intensive, especially when the reasoning process is complex. When labeled data is insufficient, the performance of models will suffer an unendurable decline. In this paper, we propose a unified framework for unsupervised complex tabular reasoning (UCTR), which generates sufficient and diverse synthetic data with complex logic for tabular reasoning tasks, assuming no human-annotated data at all. Specifically, we first utilize a random sampling strategy to collect diverse programs of different types and execute them on tables based on a "Program-Executor" module. To bridge the gap between the programs and natural language sentences, we design a powerful "NL-Generator" module to generate natural language sentences with complex logic from these programs. Since a table often occurs with its surrounding texts, we further propose novel "Table-to-Text" and "Text-to-Table" operators to handle joint table-text reasoning scenarios. This way, we can adequately exploit the unlabeled table resources to obtain a well-performed reasoning model under an unsupervised setting. Our experiments cover different tasks (question answering and fact verification) and different domains (general and specific), showing that our unsupervised methods can achieve at most 93% performance compared to supervised models. The impressive performance demonstrates that UCTR can significantly reduce the dependence on manual annotation. Moreover, we also find that it can substantially boost the supervised performance in low-resourced domains as a data augmentation technique.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question AnsweringZhenyu Li, Sunqi Fan, Yu Gu, Xiuxing Li 等AAAI 2024 · 被引用 143 次
- COMM: Concentrated Margin Maximization for Robust Document-Level Relation ExtractionZhichao Duan, Tengyu Pan, Zhenyu Li, Xiuxing Li 等AAAI 2025 · 被引用 1 次
- HCT-QA: A Benchmark for Question Answering on Human-Centric TablesMohammad Shahmeer Ahmad, Zan Ahmad Naeem, Michaël Aupetit, Ahmed K. Elmagarmid 等ICDE 2026
- ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and AnsweringYindong Zhang, Wenmian Yang, Yiquan Zhang, Weijia JiaACL 2026
相关 Paper
- Generation of Training Examples for Tabular Natural Language InferenceJean-Flavien Bussotti, Enzo Veltri, Donatello Santoro, Paolo PapottiSIGMOD 2024 · 被引用 7 次
- TaREx: Reinforcement Learning for Code-Driven Table ReasoningFangyu Lei, Jinxiang Meng, Yiming Huang, Shizhu He 等AAAI 2026
- ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning ExamplesYilun Zhao, Linyong Nan, Zhenting Qi, Rui Zhang 等EMNLP 2022 · 被引用 15 次
- CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular DataZhehao Zhang, Xitao Li, Yan Gao, Jian-Guang LouEMNLP 2023 · 被引用 3 次
- Probing How Scalable Table Data Enhances General Long-Context ReasoningHuaibing Xie, Guoliang Zhao, Yang Liu, Shihan Dou 等ICML 2026
