Lune

EMNLP2025顶会

Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuning

Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Dongmei Zhang, Surajit Chaudhuri

2025年份
2被引次数
2顶会引用

摘要

Language models such as GPT and Llama have shown remarkable ability on diverse natural language tasks, yet their performance on complex table tasks (e.g., NL-to-Code, data cleaning, etc.) continues to be suboptimal. To improve their performance, task-specific fine-tuning is often needed, which, however, require expensive human labeling and is prone to over-fitting. In this work, we propose TABLE-SPECIALIST, a self-trained fine-tuning paradigm specifically designed for table tasks. Our insight is that for each table task, there often exist two dual versions of the same task, one generative and one classification in nature. Leveraging their duality, we propose a Generator-Validator paradigm to iteratively generate-then-validate training data from language models, to finetune stronger TABLE-SPECIALIST models that can specialize in a given task, without using manually-labeled data. Extensive evaluations of TABLE-SPECIALIST on Llama, GPT-3.5 and GPT-4 suggest that our TABLE-SPECIALIST has (1) strong performance on diverse tasks over vanilla languagemodels -for example, TABLE-SPECIALIST fine-tuned on GPT-3.5 not only outperforms vanilla GPT-3.5, but can often surpass GPT-4 level quality, (2) lower cost to deploy, because when TABLE-SPECIALIST fine-tuned on GPT-3.5 achieve GPT-4 level quality, it becomes possible to deploy smaller models with lower latency/cost at comparable quality, and (3) better generalizability when evaluated across multiple benchmarks, since TABLE-SPECIALIST is fine-tuned on a broad range of training data systematically generated from diverse real tables. Our code is available at microsoft/Table-Specialist. Specialist models fine-tuned using TABLE-SPECIALIST have been integrated into Microsoft Excel for use cases such as automated data cleaning.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖