KnowProxy: Adapting Large Language Models by Knowledge-guided Proxy
Gukhyeon Lee, Yeachan Kim, Sangkeun Lee
摘要
Adapting large language models (LLMs) using smaller proxy models has been shown to improve training efficiency, where the LLMs remain frozen while the proxies are tuned on top. However, this approach typically requires access to the output probability distributions of LLMs, which are often inaccessible or unstable. To address this limitation, we propose KNOWPROXY, a knowledge-guided proxy framework in which the proxy is trained with textual knowledge rather than probability distributions. Specifically, we first elicit textual knowledge and reasoning from frozen LLMs through prompting, and then the proxy model learns to adapt this reasoning to target task distributions. We evaluate KNOWPROXY on diverse reasoning benchmarks with different fine-tuning scenarios. Comprehensive results show that KNOWPROXY achieves competitive or even better performance without direct access to probability distributions, thereby providing a scalable and versatile alternative to traditional fine-tuning. 1 * Equal contribution. 1 Our code and data are available at https://github.com/2gukhyeon/KnowProxy.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
相关 Paper
- SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMsHaoran Lou, Ziyan Liu, Chunxiao Fan, Yuexin Wu 等ICML 2026
- Task-Aware Data Selection via Proxy-Label Enhanced Distribution Matching for LLM FinetuningHao Cheng, Rui Zhang, Ling Li, Na Di 等ICLR 2026
- LightPROF: A Lightweight Reasoning Framework for Large Language Model on Knowledge GraphTu Ao, Yanhua Yu, Yuling Wang, Yang Deng 等AAAI 2025 · 被引用 28 次
- Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMsJaemin Kim, Hangeol Chang, Hyunmin Hwang, Choonghan Kim 等ICML 2026 · 被引用 1 次
- A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter SearchBaek Seong-Eun, Lee Jung-Mok, Kim Sung-Bin, Tae-Hyun OhICML 2026 · 被引用 2 次
