On the Transformation of Latent Space in Fine-Tuned NLP Models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, Firoj Alam
摘要
We study the evolution of latent space in fine-tuned NLP models. Different from the commonly used probing-framework, we opt for an unsupervised method to analyze representations. More specifically, we discover latent concepts in the representational space using hierarchical clustering. We then use an alignment function to gauge the similarity between the latent space of a pre-trained model and its fine-tuned version. We use traditional linguistic concepts to facilitate our understanding and also study how the model space transforms towards task-specific information. We perform a thorough analysis, comparing pre-trained and fine-tuned models across three models and three downstream tasks. The notable findings of our work are: i) the latent space of the higher layers evolve towards task-specific concepts, ii) whereas the lower layers retain generic concepts acquired in the pre-trained model, iii) we discovered that some concepts in the higher layers acquire polarity towards the output class, and iv) that these concepts can be used for generating adversarial triggers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Where does In-context Learning Happen in Large Language Models?Suzanna Sia, David Mueller, Kevin DuhNeurIPS 2024 · 被引用 14 次
- Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri 等EMNLP 2024 · 被引用 3 次
- Less is More: Local Intrinsic Dimensions of Contextual Language ModelsBenjamin Matthias Ruppik, Julius von Rohrscheidt, Carel van Niekerk, Michael Heck 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani 等ICLR 2022 · 被引用 74 次
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 被引用 35 次
- Asking without Telling: Exploring Latent Ontologies in Contextual RepresentationsJulian Michael, Jan A. Botha, Ian TenneyEMNLP 2020 · 被引用 3 次
- Analyzing Redundancy in Pretrained Transformer ModelsFahim Dalvi, Hassan Sajjad, Nadir Durrani, Yonatan BelinkovEMNLP 2020 · 被引用 2 次
相关 Paper
- A Closer Look at How Fine-tuning Changes BERTYichu Zhou, Vivek SrikumarACL 2022 · 被引用 84 次
- Interpretability of Language Models via Task SpacesLucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke HupkesACL 2024
- Exploring Alignment in Shared Cross-lingual SpacesBasel Mousi, Nadir Durrani, Fahim Dalvi, Majd Hawasly 等ACL 2024 · 被引用 1 次
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 被引用 46 次
- Backdoor Pre-trained Models Can Transfer to AllLujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li 等CCS 2021 · 被引用 72 次
