ST-Raptor: LLM-Powered Semi-Structured Table Question Answering
Zirui Tang, Boyu Niu, Xuanhe Zhou, Boxiu Li, Wei Zhou, Jiannan Wang, Guoliang Li, Xinyi Zhang, Fan Wu
Abstract
Semi-structured tables, widely used in real-world applications (e.g., financial reports, medical records, transactional orders), often involve flexible and complex layouts (e.g., hierarchical headers and merged cells). These tables generally rely on human analysts to interpret table layouts and answer relevant natural language questions, which is costly and inefficient. To automate the procedure, existing methods face significant challenges. First, methods like NL2SQL require converting semi-structured tables into structured ones, which often causes substantial information loss. Second, methods like NL2Code and multi-modal LLM QA struggle to understand the complex layouts of semi-structured tables and cannot accurately answer corresponding questions. To this end, we propose ST-Raptor, a tree-based framework for semi-structured table question answering ( semi-structured table QA ) using large language models. First, we introduce the Hierarchical Orthogonal Tree (HO-Tree), a structural model that captures complex semi-structured table layouts, along with an effective algorithm for constructing the tree by identifying headers, content values, and their implicit relationships. Second, we define a set of basic tree operations to guide LLMs in executing common QA tasks. Given a user question, ST-Raptor decomposes it into simpler sub-questions, generates corresponding tree operation pipelines, and conducts operation-table alignment for accurate pipeline execution. Third, we incorporate a two-stage verification mechanism: (1) forward validation checks the correctness of execution steps, while (2) backward validation evaluates answer reliability by reconstructing queries from predicted answers. To benchmark the performance, we present SSTQA, a dataset of 764 questions over 102 real-world semi-structured tables. Experiments show that ST-Raptor outperforms nine baselines by up to 20% in answer accuracy. The code is available at https://github.com/weAIDB/ST-Raptor.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f62eeff0-71ed-455e-bc18-4a9ea0c908e3Cited by top-tier papers3
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and EvaluationWei Zhou, Bolei Ma, Annemarie Friedrich, Mohsen MesgarACL 2026 · 3 citations
- Orthogonal Hierarchical Decomposition for Structure-Aware Table Understanding with Large Language ModelsBin Cao, huixian lu, chenwen ma, Ting Wang et al.ICML 2026 · 2 citations
- MoDora: Tree-Based Semi-Structured Document Analysis SystemBangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou et al.SIGMOD 2026
Builds on10
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan et al.VLDB 2024 · 165 citations
- ReAcTable: Enhancing ReAct for Table Question AnsweringYunjia Zhang, Jordan Henkel, Avrilia Floratou, Joyce Cahoon et al.VLDB 2024 · 120 citations
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu et al.ACL 2023 · 104 citations
- Decomposed Prompting: A Modular Approach for Solving Complex TasksTushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu et al.ICLR 2023 · 94 citations
- OpenSearch-SQL: Enhancing Text-to-SQL with Dynamic Few-shot and Consistency AlignmentXiangjin Xie, Guangwei Xu, Lingyan Zhao, Ruijie GuoSIGMOD 2025 · 28 citations
Related papers
- Weaver: Interweaving SQL and LLM for Table ReasoningRohit Khoja, Devanshu Gupta, Yanjie Fu, Dan Roth et al.EMNLP 2025 · 1 citation
- ASTRA: Adaptive Semantic Tree Reasoning Architecture for Complex Table Question AnsweringXiaoke Guo, Songze Li, Zhiqiang Liu, Zhaoyan Gong et al.ACL 2026 · 3 citations
- CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular TablesZhen Yang, Wei Du, Jie Wang, Wenze Zhou et al.ACL 2026
- Same Content, Different Representations: A Controlled Study for Table QAYue Zhang, Seiji Maekawa, Nikita BhutaniICLR 2026 · 5 citations
- Accurate and Regret-Aware Numerical Problem Solver for Tabular Question AnsweringYuxiang Wang, Jianzhong Qi, Junhao GanAAAI 2025 · 12 citations
