TKGT: Redefinition and A New Way of Text-to-Table Tasks Based on Real World Demands and Knowledge Graphs Augmented LLMs
Peiwen Jiang, Xinbo Lin, Zibo Zhao, Ruhui Ma, Yvonne Jie Chen, Jinhua Cheng
Abstract
The task of text-to-table receives widespread attention, yet its importance and difficulty are underestimated. Existing works use simple datasets similar to table-to-text tasks and employ methods that ignore domain structures. As a bridge between raw text and statistical analysis, the text-to-table task often deals with complex semi-structured texts that refer to specific domain topics in the real world with entities and events, especially from those of social sciences. In this paper, we analyze the limitations of benchmark datasets and methods used in the text-to-table literature and redefine the textto-table task to improve its compatibility with long text-processing tasks. Based on this redefinition, we propose a new dataset called CPL (Chinese Private Lending), which consists of judgments from China and is derived from a real-world legal academic project. We further propose TKGT (Text-KG-Table ), a two stages domain-aware pipeline, which firstly generates domain knowledge graphs (KGs) classes semiautomatically from raw text with the mixed information extraction (Mixed-IE) method, then adopts the hybrid retrieval augmented generation (Hybird-RAG) method to transform it to tables for downstream needs under the guidance of KGs classes. Experiment results show that TKGT achieves state-of-the-art (SOTA) performance on both traditional datasets and the CPL. Our data and main code are available at https://github.com/jiangpw41/TKGT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bc6c3ff-7f22-4756-a8fe-7016eb53d2c8Cited by top-tier papers3
- From Automation to Autonomy: A Survey on Large Language Models in Scientific DiscoveryTianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang et al.EMNLP 2025 · 5 citations
- TST: A Schema-Based Top-Down and Dynamic-Aware Agent of Text-to-Table TasksPeiwen Jiang, Haitong Jiang, Ruhui Ma, Yvonne Jie Chen et al.ACL 2025
- TableMix: Enhancing Multimodal Table Reasoning in MLLMs from a Data-Centric PerspectiveChaohu Liu, Shida Wang, Yubo Wang, Linli XuCVPR 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question AnsweringXiaoxin He, Yijun Tian, Yifei Sun, Nitesh V. Chawla et al.NeurIPS 2024 · 384 citations
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 211 citations
- Lift Yourself Up: Retrieval-augmented Text Generation with Self-MemoryXin Cheng, Di Luo, Xiuying Chen, Lemao Liu et al.NeurIPS 2023 · 177 citations
Related papers
- Text-to-Table: A New Way of Information ExtractionXueqing Wu, Jiacheng Zhang, Hang LiACL 2022
- ENT-DESC: Entity Description Generation by Exploring Knowledge GraphLiying Cheng, Dekun Wu, Lidong Bing, Yan Zhang et al.EMNLP 2020 · 22 citations
- Can LLMs be Good Graph Judge for Knowledge Graph Construction?Haoyu Huang, Chong Chen, Zeang Sheng, Yang Li et al.EMNLP 2025 · 7 citations
- KGPT: Knowledge-Grounded Pre-Training for Data-to-Text GenerationWenhu Chen, Yu Su, Xifeng Yan, William Yang WangEMNLP 2020 · 115 citations
- CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High QualityLiang Li, Ruiying Geng, Chengyang Fang, Bing Li et al.ACL 2023 · 2 citations
