LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder
Yi Jing, Zijun Yao, Hongzhu Guo, Lingxu Ran, Xiaozhi Wang, Lei Hou, Juanzi Li
Abstract
Large language models (LLMs) demonstrate exceptional performance on tasks requiring complex linguistic abilities, such as reference disambiguation and metaphor recognition/generation. Although LLMs possess impressive capabilities, their internal mechanisms for processing and representing linguistic knowledge remain largely opaque. Prior research on linguistic mechanisms is limited by coarse granularity, limited analysis scale, and narrow focus. In this study, we propose LinguaLens, a systematic and comprehensive framework for analyzing the linguistic mechanisms of large language models, based on Sparse Auto-Encoders (SAEs). We extract a broad set of Chinese and English linguistic features across four dimensions (morphology, syntax, semantics, and pragmatics). By employing counterfactual methods, we construct a large-scale counterfactual dataset of linguistic features for mechanism analysis. Our findings reveal intrinsic representations of linguistic knowledge in LLMs, uncover patterns of cross-layer and cross-lingual distribution, and demonstrate the potential to control model outputs. This work provides a systematic suite of resources and methods for studying linguistic mechanisms, offers strong evidence that LLMs possess genuine linguistic knowledge, and lays the foundation for more interpretable and controllable language modeling in future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ddf6302-7225-410c-928c-73fa7734a754Cited by top-tier papers2
- Understanding or Memorizing? A Case Study of German Definite Articles in Language ModelsJonathan Drechsel, Erisa Bytyqi, Steffen HerboldACL 2026 · 1 citation
- HistLens: Mapping Idea Change across Concepts and CorporaYi Jing, Weiyun Qiu, Yihang Peng, Zhifang SuiACL 2026
Builds on7
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Interpreting Language Models with Contrastive ExplanationsKayo Yin, Graham NeubigEMNLP 2022 · 32 citations
- Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language ModelsLennart Wachowiak, Dagmar GromannACL 2023 · 11 citations
- Scaling and evaluating sparse autoencodersLeo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh et al.ICLR 2025 · 10 citations
- Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation PerspectiveHanqi Yan, Yanzheng Xiang, Guangyi Chen, Yifei Wang et al.EMNLP 2024 · 1 citation
Related papers
- Toward Faithful Retrieval-Augmented Generation with Sparse AutoencodersGuangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha et al.ICLR 2026 · 8 citations
- Metaphors in Pre-Trained Language Models: Probing and Generalization Across Datasets and LanguagesEhsan Aghazadeh, Mohsen Fayyaz, Yadollah YaghoobzadehACL 2022
- The Same but Different: Structural Similarities and Differences in Multilingual Language ModelingRuochen Zhang, Qinan Yu, Matianyu Zang, Carsten Eickhoff et al.ICLR 2025
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsHaoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li et al.IEEE VIS 2025 · 3 citations
- Large Multi-modal Models Can Interpret Features in Large Multi-modal ModelsKaichen Zhang, Yifei Shen, Bo Li, Ziwei LiuICCV 2025 · 2 citations
