AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language models
José Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari A. Sahraoui
摘要
The objective of pre-trained language models is to learn contextual representations of textual data. Pre-trained language models have become mainstream in natural language processing and code modeling. Using probes, a technique to study the linguistic properties of hidden vector spaces, previous works have shown that these pre-trained language models encode simple linguistic properties in their hidden representations. However, none of the previous work assessed whether these models encode the whole grammatical structure of a programming language. In this paper, we prove the existence of a syntactic subspace, lying in the hidden representations of pre-trained language models, which contain the syntactic information of the programming language. We show that this subspace can be extracted from the models’ representations and define a novel probing method, the AST-Probe, that enables recovering the whole abstract syntax tree (AST) of an input code snippet. In our experimentations, we show that this syntactic subspace exists in five state-of-the-art pre-trained language models. In addition, we highlight that the middle layers of the models are the ones that encode most of the AST information. Finally, we estimate the optimal size of this syntactic subspace and show that its dimension is substantially lower than those of the models’ representation spaces. This suggests that pre-trained language models use a small part of their representation spaces to encode syntactic information of the programming languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Towards Efficient Fine-Tuning of Pre-trained Code Models: An Experimental Study and BeyondEnsheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du 等ISSTA 2023 · 被引用 37 次
- Better Context Makes Better Code Language Models: A Case Study on Function Call Argument CompletionHengzhi Pei, Jinman Zhao, Leonard Lausen, Sheng Zha 等AAAI 2023 · 被引用 30 次
- A Learning-Based Approach to Static Program SlicingAashish Yadavally, Yi Li, Shaohua Wang, Tien N. NguyenOOPSLA 2024 · 被引用 15 次
- The EarlyBIRD Catches the Bug: On Exploiting Early Layers of Encoder Models for More Efficient Code ClassificationAnastasiia Grishina, Max Hort, Leon MoonenFSE 2023 · 被引用 15 次
- The Devil is in the Tails: How Long-Tailed Code Distributions Impact Large Language ModelsXin Zhou, Kisub Kim, Bowen Xu, Jiakun Liu 等ASE 2023 · 被引用 10 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 被引用 98 次
- What Do They Capture? - A Structural Analysis of Pre-Trained Language Models for Source CodeYao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui 等ICSE 2022 · 被引用 66 次
相关 Paper
- AST-T5: Structure-Aware Pretraining for Code Generation and UnderstandingLinyuan Gong, Mostafa Elhoushi, Alvin CheungICML 2024 · 被引用 42 次
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 被引用 5 次
- A Polar coordinate system represents syntax in large language modelsPablo Diego-Simón, Stéphane d'Ascoli, Emmanuel Chemla, Yair Lakretz 等NeurIPS 2024 · 被引用 27 次
- Language-Agnostic Representation Learning of Source Code from Structure and ContextDaniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec 等ICLR 2021 · 被引用 131 次
- A Latent-Variable Model for Intrinsic ProbingKarolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell 等AAAI 2023 · 被引用 6 次
