Jellyfish: Instruction-Tuning Local Large Language Models for Data Preprocessing
Haochen Zhang, Yuyang Dong, Chuan Xiao, Masafumi Oyamada
摘要
This paper explores the utilization of LLMs for data preprocessing (DP), a crucial step in the data mining pipeline that transforms raw data into a clean format conducive to easy processing. Whereas the use of LLMs has sparked interest in devising universal solutions to DP, recent initiatives in this domain typically rely on GPT APIs, raising inevitable data breach concerns. Unlike these approaches, we consider instruction-tuning local LLMs (7 -13B models) as universal DP task solvers that operate on a local, single, and low-priced GPU, ensuring data security and enabling further customization. We select a collection of datasets across four representative DP tasks and construct instruction tuning data using data configuration, knowledge injection, and reasoning data distillation techniques tailored to DP. By tuning Mistral-7B, Llama 3-8B, and OpenOrca-Platypus2-13B, our models, namely, Jellyfish-7B/8B/13B, deliver competitiveness compared to GPT-3.5/4 models and strong generalizability to unseen tasks while barely compromising the base models' abilities in NLP tasks. Meanwhile, Jellyfish offers enhanced reasoning capabilities compared to GPT-3.5.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- On LLM-Enhanced Mixed-Type Data Imputation with High-Order Message PassingJianmin Wang, Kai Wang, Ying Zhang, Wenjie Zhang 等VLDB 2025 · 被引用 15 次
- QUEST: Query Optimization in Unstructured Document AnalysisZhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin 等VLDB 2025 · 被引用 9 次
- Empowering Tabular Data Preparation with Language Models: Why and How?Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang 等ACL 2026 · 被引用 4 次
- BEACON: Budget-Aware Entity Matching Across DomainsNicholas Pulsone, Roee Shraga, Gregory GorenSIGMOD 2026 · 被引用 2 次
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing 等VLDB 2026
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- KnowTrans: Boosting Transferability of Data Preparation LLMs via Knowledge AugmentationYuhang Ge, Fengyu Li, Yuren Mao, Yanbo Yang 等ICDE 2025
- MAmmoTH2: Scaling Instructions from the WebXiang Yue, Tianyu Zheng, Ge Zhang, Wenhu ChenNeurIPS 2024 · 被引用 176 次
- Efficient Mixture of Experts based on Large Language Models for Low-Resource Data PreprocessingMengyi Yan, Yaoshu Wang, Kehan Pang, Min Xie 等KDD 2024 · 被引用 9 次
- LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuningWei Huang, Anda Cheng, Yinggui Wang, Lei Wang 等VLDB 2026 · 被引用 1 次
- Making Large Language Models Better Data CreatorsDong-Ho Lee, Jay Pujara, Mohit Sewak, Ryen White 等EMNLP 2023 · 被引用 13 次
