NormXLogit: The Head-on-Top Never Lies
Sina Abbasi, Mohammad Reza Modarres, Mohammad Taher Pilehvar
Abstract
With new large language models (LLMs) emerging frequently, it is important to consider the potential value of model-agnostic approaches that can provide interpretability across a variety of architectures. While recent advances in LLM interpretability show promise, many rely on complex, model-specific methods with high computational costs. To address these limitations, we propose Nor-mXLogit, a novel technique for assessing the significance of individual input tokens. This method operates based on the input and output representations associated with each token. First, we demonstrate that the norm of word embeddings can be utilized as a measure of token importance. Second, we reveal a significant relationship between a token's importance and how predictive its representation is of the model's final output. Extensive analyses indicate that our approach outperforms existing gradient-based methods in terms of faithfulness and offers competitive performance compared to leading architecture-specific techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28f1e201-b12e-4336-9c54-242f656c145bCited by top-tier papers1
Ask how each one uses itBuilds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
- Interpreting Language Models with Contrastive ExplanationsKayo Yin, Graham NeubigEMNLP 2022 · 32 citations
- Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2021 · 28 citations
Related papers
- Faithfulness Measurable Masked Language ModelsAndreas Madsen, Siva Reddy, Sarath ChandarICML 2024 · 6 citations
- Enhancing Automated Interpretability with Output-Centric Feature DescriptionsYoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger et al.ACL 2025
- Rethinking Layer Relevance in Large Language Models Beyond Cosine SimilarityCristian Hinostroza, Rodrigo Toro Icarte, Christ Devia, Andres Carvallo et al.ICLR 2026 · 4 citations
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha et al.ICML 2026 · 6 citations
- Forget What Matters, Keep the Rest: Selective Unlearning of Informative TokensSeunghee Koh, Sunghyun Baek, Youngdong Kim, Junmo KimACL 2026 · 1 citation
