SADA: Bridging In-Context Learning and Fine-Tuning via State-Aligned Distillation Adapters
Wenhao Gao, Tianlong Wang, Wei Jia, Linhao Zhang, Aiwei Liu, Miao Fan, Xiao Zhou
Abstract
Prompt-based in-context learning (ICL) and parameter fine-tuning are two dominant paradigms for incorporating external information into large language models (LLMs), but they incur high inference costs or require expensive retraining. To bridge this gap, context-to-parameter mapping converts prompts into temporary adapter weights. However, we identify a critical failure mode in existing methods: hiddenstate collapse, where the adapter-augmented model's internal states diverge sharply from the full-context oracle in deeper layers. We trace this failure to two coupled gaps: suboptimal Input-Selection and inadequate Supervision-Signal. To address these issues, we propose SADA (State-Aligned Distillation Adapters). We establish the attention-block output as a principled feature interface to improve input selection and introduce statealignment distillation to enforce consistency between the adapter-augmented model and the full-context oracle. Experiments on long-context language modeling (PG19) and downstream NLU and summarization benchmarks show that SADA consistently outperforms strong baselines like StreamAdapter and GenerativeAdapter, achieving performance comparable to ICL while significantly reducing memory footprint and latency. We further analyze when parameterized context compression is effective and when explicit context retention remains preferable. Our code is available at https://github. com/Taylor-Gavel/SADA.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 063405cb-86f1-4cc0-bff0-5c4352f08f2bBuilds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye et al.ACL 2026 · 24 citations
- Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward PassTong Chen, Hao Fang, Patrick Xia, Xiaodong Liu et al.ICLR 2025
- Doc-to-LoRA: Learning to Instantly Internalize ContextsRujikorn Charakorn, Edoardo Cetin, Shinnosuke Uesaka, Robert LangeICML 2026 · 27 citations
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsZhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao et al.NeurIPS 2025 · 22 citations
- Context Tuning for In-Context OptimizationJack Lu, Ryan Teehan, Zhenbang Yang, Mengye RenICML 2026
