The POLAR Framework: Polar Opposites Enable Interpretability of Pre-Trained Word Embeddings
Binny Mathew, Sandipan Sikdar, Florian Lemmerich, Markus Strohmaier
摘要
We introduce 'POLAR' -a framework that adds interpretability to pre-trained word embeddings via the adoption of semantic differentials. Semantic differentials are a psychometric construct for measuring the semantics of a word by analysing its position on a scale between two polar opposites (e.g., cold -hot, soft -hard). The core idea of our approach is to transform existing, pre-trained word embeddings via semantic differentials to a new "polar" space with interpretable dimensions defined by such polar opposites. Our framework also allows for selecting the most discriminative dimensions from a set of polar dimensions provided by an oracle, i.e., an external source. We demonstrate the effectiveness of our framework by deploying it to various downstream tasks, in which our interpretable word embeddings achieve a performance that is comparable to the original word embeddings. We also show that the interpretable dimensions selected by our framework align with human judgement. Together, these results demonstrate that interpretability can be added to word embeddings without compromising performance. Our work is relevant for researchers and engineers interested in interpreting pre-trained word embeddings. CCS CONCEPTS • Computing methodologies → Machine learning approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Discovering Differences in the Representation of People using Contextualized Semantic AxesLi Lucy, Divya Tadimeti, David BammanEMNLP 2022 · 被引用 7 次
- Discovering Universal Geometry in Embeddings with ICAHiroaki Yamagiwa, Momose Oyama, Hidetoshi ShimodairaEMNLP 2023 · 被引用 6 次
- Interpreting Embedding Spaces by ConceptualizationAdi Simhi, Shaul MarkovitchEMNLP 2023 · 被引用 6 次
- Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content ModelKathleen C. Fraser, Isar Nejadgholi, Svetlana KiritchenkoACL 2021
- Interpretable Embeddings with Sparse Autoencoders: A Data Analysis ToolkitNick Jiang, Xiaoqing Sun, Lisa Dunlap, Lewis Smith 等ICML 2026
相关 Paper
- Interpretable Debiasing of Vectorized Language Representations with Iterative OrthogonalizationPrince Osei Aboagye, Yan Zheng, Jack Shunn, Chin-Chia Michael Yeh 等ICLR 2023
- Introducing Orthogonal Constraint in Structural ProbesTomasz Limisiewicz, David MarecekACL 2021
- PRISM: A Framework for Producing Interpretable Political Bias Embeddings with Political-Aware Cross-EncoderYiqun Sun, Qiang Huang, Anthony Kum Hoe Tung, Jun YuACL 2025 · 被引用 2 次
- Adaptive Probabilistic Word EmbeddingShuangyin Li, Yu Zhang, Rong Pan, Kaixiang MoWWW 2020 · 被引用 10 次
- A Method for Studying Semantic Construal in Grammatical Constructions with Interpretable Contextual Embedding SpacesGabriella Chronis, Kyle Mahowald, Katrin ErkACL 2023 · 被引用 3 次
