HiFi: High-Information Attention Heads Hold for Parameter-Efficient Model Adaptation
Anchun Gui, Han Xiao
摘要
To fully leverage the advantages of large-scale pre-trained language models (PLMs) on downstream tasks, it has become a ubiquitous adaptation paradigm to fine-tune the entire parameters of PLMs. However, this paradigm poses issues of inefficient updating and resource overconsuming for fine-tuning in data-scarce and resource-limited scenarios, because of the large scale of parameters in PLMs. To alleviate these concerns, in this paper, we propose a parameterefficient fine-tuning method HiFi, that is, only the highly informative and strongly correlated attention heads for the specific task are finetuned. To search for those significant attention heads, we develop a novel framework to analyze the effectiveness of heads. Specifically, we first model the relationship between heads into a graph from two perspectives of information richness and correlation, and then apply PageRank algorithm to determine the relative importance of each head. Extensive experiments on the GLUE benchmark demonstrate the effectiveness of our method, and show that HiFi obtains state-of-the-art performance over the prior baselines. * Corresponding author. (a) Full Fine-tuning (b1) Adapter-like (c) Non-structured Method (b2) HiFi (Ours) (b) Structured Method Updated Param. Extra Updated Param.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando 等ICML 2024 · 被引用 8 次
- Prototype-based HyperAdapter for Sample-Efficient Multi-task TuningHao Zhao, Jie Fu, Zhaofeng HeEMNLP 2023 · 被引用 3 次
- Steering at the Source: Style Modulation Heads for Robust Persona ControlYoshihiro Izawa, Gouki Minegishi, Koshi Eguchi, Sosuke Hosokawa 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuningRunxin Xu, Fuli Luo, Zhiyuan Zhang, Chuanqi Tan 等EMNLP 2021 · 被引用 129 次
相关 Paper
- From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM AlignmentHao Chen, Qi Zhang, Liyao Li, Zhanming Shen 等ICML 2026
- APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language ModelsQifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu 等EMNLP 2023 · 被引用 25 次
- Parameter-Efficient Fine-Tuning without Introducing New LatencyBaohao Liao, Yan Meng, Christof MonzACL 2023 · 被引用 26 次
- HeadMap: Locating and Enhancing Knowledge Circuits in LLMsXuehao Wang, Liyuan Wang, Binghuai Lin, Yu ZhangICLR 2025
- GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream AdaptationSungmin Kang, Jisoo Kim, Salman Avestimehr, Sunwoo LeeAAAI 2026
