Weak-to-Strong Prompts with Lightweight-to-Powerful LLMs for High-Accuracy, Low-Cost, and Explainable Data Transformation
Changlun Li, Chenyu Yang, Yuyu Luo, Ju Fan, Nan Tang
Abstract
Data transformation poses significant challenges due to the wide diversity in input data formats and different requirements. Existing approaches—including human-driven, algorithmic, and large language model (LLM)-based solutions—each exhibits trade-offs in terms of cost, accuracy, and the range of supported transformations. To address these limitations, we propose MegaTran , a novel framework for generating accurate and cost-effective data transformation code. MegaTran employs a two-stage process: Weak2StrongPrompt , which converts a user's weak prompt (a loosely specified user input) into a strong, structured prompt, and Prompt2Code , which generates transformation code based on this refined prompt. In Weak2StrongPrompt , a fine-tuned lightweight LLM predicts the transformation type and generates a detailed task description from the user's input. In Prompt2Code , a powerful LLM generates the corresponding transformation code, guided by two key optimizations: (1) Sanity-check Reflection with checklist , which iteratively debugs and refines the code by addressing errors; and (2) Lazy-RAG , a retrieval-augmented generation technique that retrieves relevant code snippets or documentation from external resources ( e.g. , GitHub, DataPrep) to enhance code quality. Extensive experiments show that MegaTran achieves results varying from +2.2% to +26.1% accuracy improvement compared with SoTA methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32c35c0b-e08a-4621-afba-6d1f359dae1dCited by top-tier papers3
- DeepPrep: An LLM-Powered Agentic System for Autonomous Data PreparationMeihao Fan, Ju Fan, Yuxin Zhang, Shaolei Zhang et al.VLDB 2026 · 4 citations
- TuneAhead: Predicting Fine-tuning Performance Before Training BeginsYuxiang Luo, Haonan Long, Chen Wang, Qiqi Duan et al.ICML 2026
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
Related papers
- LOOPRAG: Enhancing Loop Transformation Optimization with Retrieval-Augmented Large Language ModelsYijie Zhi, Yayu Cao, Jianhua Dai, Xiaoyang Han et al.ASPLOS 2026 · 1 citation
- ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User PromptsJiajun Zhu, Xinyu Cheng, Zhongsu Luo, Yunfan Zhou et al.UIST 2025 · 1 citation
- QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language ModelsQirui Zhou, Yuanbo Wen, Ruizhi Chen, Ke Gao et al.AAAI 2025 · 7 citations
- Reliable and Cost-Effective Exploratory Data Analysis via Graph-Guided RAGMossad Helali, Yutai Luo, Tae Jun Ham, Jim Plotts et al.EMNLP 2025
- Preference-Guided Refactored Tuning for Retrieval Augmented Code GenerationXinyu Gao, Yun Xiong, Deze Wang, Zhenhan Guan et al.ASE 2024 · 1 citation
