Realistic Training Data Generation and Rule Enhanced Decoding in LLM for NameGuess
Yikuan Xia, Jiazun Chen, Sujian Li, Jun Gao
Abstract
The wide use of abbreviated column names (derived from English words or Chinese Pinyin) in database tables poses significant challenges for table-centric tasks in natural language processing and database management. Such a column name expansion task, referred to as the NameGuess task, has previously been addressed by fine-tuning Large Language Models (LLMs) on synthetically generated rule-based data. However, the current approaches yield suboptimal performance due to two fundamental limitations: 1) the rule-generated abbreviation data fails to reflect real-world distribution, and 2) the failure of LLMs to follow the rulesensitive patterns in NameGuess persistently. For the data realism issue, we propose a novel approach that integrates a subsequence abbreviation generator trained on human-annotated data and collects non-subsequence abbreviations to improve the training set. For the rule violation issue, we propose a decoding system constrained on an automaton that represents the rules of abbreviation expansion. We extended the original English NameGuess test set to include non-subsequence and PinYin scenarios. Experimental results show that properly tuned 7/8B moderate-size LLMs with a refined decoding system can surpass the few-shot performance of state-of-the-art LLMs, such as the GPT-4 series. The code and data are presented in the supplementary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 461ab2a6-0bf5-4365-8b0f-773631c5da26Builds on8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 200 citations
- Valentine: Evaluating Matching Techniques for Dataset DiscoveryChristos Koutras, George Siachamis, Andra Ionescu, Kyriakos Psarakis et al.ICDE 2021 · 87 citations
Related papers
- NameGuess: Column Name Expansion for Tabular DataJiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan, Shen Wang et al.EMNLP 2023 · 6 citations
- Exploring and Adapting Chinese GPT to Pinyin Input MethodMinghuan Tan, Yong Dai, Duyu Tang, Zhangyin Feng et al.ACL 2022 · 13 citations
- ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQLYang Qin, Chao Chen, Zhihang Fu, Ze Chen et al.ICLR 2025
- Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuningJunjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong et al.EMNLP 2025 · 2 citations
- SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingTianyu Yu, Chengyue Jiang, Chao Lou, Shen Huang et al.AAAI 2024 · 30 citations
