Introducing Orthogonal Constraint in Structural Probes
Tomasz Limisiewicz, David Marecek
Abstract
With the recent success of pre-trained models in NLP, a significant focus was put on interpreting their representations. One of the most prominent approaches is structural probing (Hewitt and Manning, 2019) , where a linear projection of word embeddings is performed in order to approximate the topology of dependency structures. In this work, we introduce a new type of structural probing, where the linear projection is decomposed into 1. isomorphic space rotation; 2. linear scaling that identifies and scales the most relevant dimensions. In addition to syntactic dependency, we evaluate our method on novel tasks (lexical hypernymy and position in a sentence). We jointly train the probes for multiple tasks and experimentally show that lexical and syntactic information is separated in the representations. Moreover, the orthogonal constraint makes the Structural Probes less vulnerable to memorization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- A Polar coordinate system represents syntax in large language modelsPablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz et al.NeurIPS 2024 · 27 citations
- AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language modelsJosé Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari A. SahraouiASE 2022 · 18 citations
- Probing for Labeled Dependency TreesMax Müller-Eberstein, Rob van der Goot, Barbara PlankACL 2022 · 10 citations
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 5 citations
- Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?Ahmed Alajrami, Katerina Margatina, Nikolaos AletrasEMNLP 2023 · 1 citation
Builds on4
- Are All Good Word Vector Spaces Isomorphic?Ivan Vulic, Sebastian Ruder, Anders SøgaardEMNLP 2020 · 7 citations
- Pareto Probing: Trading Off Accuracy for ComplexityTiago Pimentel, Naomi Saphra, Adina Williams, Ryan CotterellEMNLP 2020 · 6 citations
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 3 citations
- Do Neural Language Models Show Preferences for Syntactic Formalisms?Artur Kulmizev, Vinit Ravishankar, Mostafa Abdou, Joakim NivreACL 2020 · 1 citation
Related papers
- Probing BERT in Hyperbolic SpacesBoli Chen, Yao Fu, Guangwei Xu, Pengjun Xie et al.ICLR 2021 · 19 citations
- A Latent-Variable Model for Intrinsic ProbingKarolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell et al.AAAI 2023 · 6 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- A Closer Look at How Fine-tuning Changes BERTYichu Zhou, Vivek SrikumarACL 2022 · 84 citations
- Probing Linguistic Features of Sentence-Level Representations in Relation ExtractionChristoph Alt, Aleksandra Gabryszak, Leonhard HennigACL 2020 · 29 citations
