AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language models
José Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari A. Sahraoui
Abstract
The objective of pre-trained language models is to learn contextual representations of textual data. Pre-trained language models have become mainstream in natural language processing and code modeling. Using probes, a technique to study the linguistic properties of hidden vector spaces, previous works have shown that these pre-trained language models encode simple linguistic properties in their hidden representations. However, none of the previous work assessed whether these models encode the whole grammatical structure of a programming language. In this paper, we prove the existence of a syntactic subspace, lying in the hidden representations of pre-trained language models, which contain the syntactic information of the programming language. We show that this subspace can be extracted from the models’ representations and define a novel probing method, the AST-Probe, that enables recovering the whole abstract syntax tree (AST) of an input code snippet. In our experimentations, we show that this syntactic subspace exists in five state-of-the-art pre-trained language models. In addition, we highlight that the middle layers of the models are the ones that encode most of the AST information. Finally, we estimate the optimal size of this syntactic subspace and show that its dimension is substantially lower than those of the models’ representation spaces. This suggests that pre-trained language models use a small part of their representation spaces to encode syntactic information of the programming languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ee696f1-7f49-4f77-97c8-933bc942d1b2Cited by top-tier papers7
- Towards Efficient Fine-Tuning of Pre-trained Code Models: An Experimental Study and BeyondEnsheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du et al.ISSTA 2023 · 37 citations
- Better Context Makes Better Code Language Models: A Case Study on Function Call Argument CompletionHengzhi Pei, Jinman Zhao, Leonard Lausen, Sheng Zha et al.AAAI 2023 · 30 citations
- A Learning-Based Approach to Static Program SlicingAashish Yadavally, Yi Li, Shaohua Wang, Tien N. NguyenOOPSLA 2024 · 15 citations
- The EarlyBIRD Catches the Bug: On Exploiting Early Layers of Encoder Models for More Efficient Code ClassificationAnastasiia Grishina, Max Hort, Leon MoonenFSE 2023 · 15 citations
- The Devil is in the Tails: How Long-Tailed Code Distributions Impact Large Language ModelsXin Zhou, Kisub Kim, Bowen Xu, Jiakun Liu et al.ASE 2023 · 10 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 98 citations
- What Do They Capture? - A Structural Analysis of Pre-Trained Language Models for Source CodeYao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui et al.ICSE 2022 · 66 citations
Related papers
- AST-T5: Structure-Aware Pretraining for Code Generation and UnderstandingLinyuan Gong, Mostafa Elhoushi, Alvin CheungICML 2024 · 42 citations
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 5 citations
- A Polar coordinate system represents syntax in large language modelsPablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz et al.NeurIPS 2024 · 27 citations
- Language-Agnostic Representation Learning of Source Code from Structure and ContextDaniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec et al.ICLR 2021 · 131 citations
- A Latent-Variable Model for Intrinsic ProbingKarolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell et al.AAAI 2023 · 6 citations
