On the Transformation of Latent Space in Fine-Tuned NLP Models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, Firoj Alam
Abstract
We study the evolution of latent space in fine-tuned NLP models. Different from the commonly used probing-framework, we opt for an unsupervised method to analyze representations. More specifically, we discover latent concepts in the representational space using hierarchical clustering. We then use an alignment function to gauge the similarity between the latent space of a pre-trained model and its fine-tuned version. We use traditional linguistic concepts to facilitate our understanding and also study how the model space transforms towards task-specific information. We perform a thorough analysis, comparing pre-trained and fine-tuned models across three models and three downstream tasks. The notable findings of our work are: i) the latent space of the higher layers evolve towards task-specific concepts, ii) whereas the lower layers retain generic concepts acquired in the pre-trained model, iii) we discovered that some concepts in the higher layers acquire polarity towards the output class, and iv) that these concepts can be used for generating adversarial triggers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ebef1cb-a449-4c89-b1ab-fedb747285d8Cited by top-tier papers3
- Where does In-context Learning Happen in Large Language Models?Suzanna Sia, David Mueller, Kevin DuhNeurIPS 2024 · 14 citations
- Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri et al.EMNLP 2024 · 3 citations
- Less is More: Local Intrinsic Dimensions of Contextual Language ModelsBenjamin Matthias Ruppik, Julius von Rohrscheidt, Carel van Niekerk, Michael Heck et al.NeurIPS 2025 · 1 citation
Builds on5
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani et al.ICLR 2022 · 74 citations
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 35 citations
- Asking without Telling: Exploring Latent Ontologies in Contextual RepresentationsJulian Michael, Jan A. Botha, Ian TenneyEMNLP 2020 · 3 citations
- Analyzing Redundancy in Pretrained Transformer ModelsFahim Dalvi, Hassan Sajjad, Nadir Durrani, Yonatan BelinkovEMNLP 2020 · 2 citations
Related papers
- A Closer Look at How Fine-tuning Changes BERTYichu Zhou, Vivek SrikumarACL 2022 · 84 citations
- Interpretability of Language Models via Task SpacesLucas Weber, Jaap Jumelet, Elia Bruni, Dieuwke HupkesACL 2024
- Exploring Alignment in Shared Cross-lingual SpacesBasel Mousi, Nadir Durrani, Fahim Dalvi, Majd Hawasly et al.ACL 2024 · 1 citation
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 46 citations
- Backdoor Pre-trained Models Can Transfer to AllLujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li et al.CCS 2021 · 72 citations
