Lune

ACL2023顶会

HiFi: High-Information Attention Heads Hold for Parameter-Efficient Model Adaptation

Anchun Gui, Han Xiao

2023年份
1被引次数
3顶会引用

摘要

To fully leverage the advantages of large-scale pre-trained language models (PLMs) on downstream tasks, it has become a ubiquitous adaptation paradigm to fine-tune the entire parameters of PLMs. However, this paradigm poses issues of inefficient updating and resource overconsuming for fine-tuning in data-scarce and resource-limited scenarios, because of the large scale of parameters in PLMs. To alleviate these concerns, in this paper, we propose a parameterefficient fine-tuning method HiFi, that is, only the highly informative and strongly correlated attention heads for the specific task are finetuned. To search for those significant attention heads, we develop a novel framework to analyze the effectiveness of heads. Specifically, we first model the relationship between heads into a graph from two perspectives of information richness and correlation, and then apply PageRank algorithm to determine the relative importance of each head. Extensive experiments on the GLUE benchmark demonstrate the effectiveness of our method, and show that HiFi obtains state-of-the-art performance over the prior baselines. * Corresponding author. (a) Full Fine-tuning (b1) Adapter-like (c) Non-structured Method (b2) HiFi (Ours) (b) Structured Method Updated Param. Extra Updated Param.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 81f11eed-b1e9-4d7b-ae8b-be8aec9d3dbf

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖