Analyzing the Rapid Generalization of SFT via the Perspective of Attention Head Activation Patterns
Yang Zhao, Li Du, Xiao Ding, Kai Xiong, Ting Liu, Bing Qin
Abstract
LLMs' performance on complex tasks is still unsatisfactory. A key issue is that presently LLMs learn in a data-driven schema, while the instructions about these complex tasks are both scarce and hard to collect or construct. On the contrary, a prominent phenomenon is that LLMs can learn rather fast on simpler tasks with adequate prior knowledge captured during pretraining stage. Thus, if the prerequisite and mechanism of such rapid generalization could be elucidated, it could enhance the efficiency and effectiveness of the LLM's ability to learn complex tasks. Thus, in this paper, we employ a gradient-based method, to dissect the process that the SFT process adapts LLMs to downstream tasks via the perspective of attention patterns. We find that: (1) LLMs selectively activate task-specific attention heads during SFT; (2) activation patterns for complex tasks are combinations of basic task patterns; and (3) changes in a few parameters can significantly impact activation patterns after SFT on a small number of samples. Based on these insights, experiments are conducted to actually enhance the efficiency and effectiveness of SFT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40b964e8-ddf5-4e3e-8af0-9827ff91bcb0Builds on7
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang et al.ICML 2024 · 605 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
- Self-Attention Attribution: Interpreting Information Interactions Inside TransformerYaru Hao, Li Dong, Furu Wei, Ke XuAAAI 2021 · 282 citations
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam et al.AAAI 2023 · 234 citations
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data CompositionGuanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li et al.ACL 2024 · 39 citations
Related papers
- Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head MaskingSenyu Han, Hongchuan Zeng, Kai Yu, Lu ChenICML 2025
- Interpretability of Language Models via Task SpacesLucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke HupkesACL 2024
- Shared Lexical Task Representations Explain Behavioral Variability In LLMsZhuonan Yang, Jacob Xiaochen Li, Francisco Velez, Eric Todd et al.ICML 2026
- Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language ModelsZheng Zhao, Yftah Ziser, Shay B. CohenEMNLP 2024
- Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data SelectionYang Zhao, Li Du, Xiao Ding, Yangou Ouyang et al.ACL 2025 · 3 citations
