On Affine Homotopy between Language Encoders
Robin Chan, Reda Boumasmoud, Anej Svete, Yuxin Ren, Qipeng Guo, Zhijing Jin, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, Mennatallah El-Assady, Ryan Cotterell
摘要
Pre-trained language encoders -- functions that represent text as vectors -- are an integral component of many NLP tasks. We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar? We contend that a faithful measure of similarity needs to be intrinsic, that is, task-independent, yet still be informative of extrinsic similarity -- the performance on downstream tasks. It is common to consider two encoders similar if they are homotopic, i.e., if they can be aligned through some transformation. In this spirit, we study the properties of affine alignment of language encoders and its implications on extrinsic similarity. We find that while affine alignment is fundamentally an asymmetric notion of similarity, it is still informative of extrinsic similarity. We confirm this on datasets of natural language representations. Beyond providing useful bounds on extrinsic similarity, affine intrinsic similarity also allows us to begin uncovering the structure of the space of pre-trained encoders by defining an order over them.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- On the Reasoning Abilities of Masked Diffusion Language ModelsAnej Svete, Ashish SabharwalICLR 2026 · 被引用 8 次
- Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don'tAnej Svete, William Merrill, Ryan Cotterell, Ashish SabharwalICML 2026 · 被引用 3 次
- Tracking Equivalent Mechanistic Interpretations Across Neural NetworksAlan Sun, Mariya TonevaICLR 2026 · 被引用 1 次
- On the Representational Capacity of Neural Language Models with Chain-of-Thought ReasoningFranz Nowak, Anej Svete, Alexandra Butoi, Ryan CotterellACL 2024
- Gumbel Counterfactual Generation From Language ModelsShauli Ravfogel, Anej Svete, Vésteinn Snæbjarnarson, Ryan CotterellICLR 2025
它引用的顶会 Paper7
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Revisiting Model Stitching to Compare Neural RepresentationsYamini Bansal, Preetum Nakkiran, Boaz BarakNeurIPS 2021 · 被引用 253 次
- Generalized Shape Metrics on Neural RepresentationsAlex H. Williams, Erin Kunz, Simon Kornblith, Scott W. LindermanNeurIPS 2021 · 被引用 182 次
- The MultiBERTs: BERT Reproductions for Robustness AnalysisThibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei 等ICLR 2022 · 被引用 106 次
- Grounding Representation Similarity Through Statistical TestingFrances Ding, Jean-Stanislas Denain, Jacob SteinhardtNeurIPS 2021 · 被引用 88 次
相关 Paper
- Latent Space Translation via Semantic AlignmentValentino Maiorca, Luca Moschella, Antonio Norelli, Marco Fumero 等NeurIPS 2023 · 被引用 59 次
- Connecting Neural Models Latent Geometries with Relative Geodesic RepresentationsHanlin Yu, Berfin Inal, Georgios Arvanitidis, Søren Hauberg 等NeurIPS 2025 · 被引用 5 次
- A Latent-Variable Model for Intrinsic ProbingKarolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell 等AAAI 2023 · 被引用 6 次
- Connecting Pre-trained Language Model and Downstream Task via Properties of RepresentationChenwei Wu, Holden Lee, Rong GeNeurIPS 2023 · 被引用 8 次
- Pretraining with Artificial Language: Studying Transferable Knowledge in Language ModelsRyokan Ri, Yoshimasa TsuruokaACL 2022 · 被引用 40 次
