What Do They Capture? - A Structural Analysis of Pre-Trained Language Models for Source Code
Yao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui, Guandong Xu, Hai Jin
摘要
Recently, many pre-trained language models for source code have been proposed to model the context of code and serve as a basis for downstream code intelligence tasks such as code completion, code search, and code summarization. These models leverage masked pre-training and Transformer and have achieved promising results. However, currently there is still little progress regarding interpretability of existing pre-trained code models. It is not clear why these models work and what feature correlations they can capture. In this paper, we conduct a thorough structural analysis aiming to provide an interpretation of pre-trained language models for source code (e.g., CodeBERT, and GraphCodeBERT) from three distinctive perspectives: (1) attention analysis, (2) probing on the word embedding, and (3) syntax tree induction. Through comprehensive analysis, this paper reveals several insightful findings that may inspire future studies: (1) Attention aligns strongly with the syntax structure of code. (2) Pre-training language models of code can preserve the syntax structure of code in the intermediate representations of each Transformer layer. (3) The pre-trained models of code have the ability of inducing syntax trees of code. Theses findings suggest that it may be helpful to incorporate the syntax structure of code into the process of pre-training for better code representations. CCS CONCEPTS • Software and its engineering → Reusability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Prompt-tuned Code Language Model as a Neural Knowledge Base for Type Inference in Statically-Typed Partial CodeQing Huang, Zhiqiang Yuan, Zhenchang Xing, Xiwei Xu 等ASE 2022 · 被引用 40 次
- Diet code is healthy: simplifying programs for pre-trained models of codeZhaowei Zhang, Hongyu Zhang, Beijun Shen, Xiaodong GuFSE 2022 · 被引用 39 次
- Towards Efficient Fine-Tuning of Pre-trained Code Models: An Experimental Study and BeyondEnsheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du 等ISSTA 2023 · 被引用 37 次
- Graph Neural Networks for Vulnerability Detection: A Counterfactual ExplanationZhaoyang Chu, Yao Wan, Qian Li, Yang Wu 等ISSTA 2024 · 被引用 19 次
- AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language modelsJosé Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, Houari A. SahraouiASE 2022 · 被引用 18 次
它引用的顶会 Paper8
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun 等ICSE 2020 · 被引用 242 次
相关 Paper
- GrammarT5: Grammar-Integrated Pretrained Encoder-Decoder Neural Model for CodeQihao Zhu, Qingyuan Liang, Zeyu Sun, Yingfei Xiong 等ICSE 2024 · 被引用 10 次
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng 等FSE 2022 · 被引用 148 次
- An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program RepairKai Huang, Xiangxin Meng, Jian Zhang, Yang Liu 等ASE 2023 · 被引用 91 次
- Natural Is the Best: Model-Agnostic Code Simplification for Pre-trained Large Language ModelsYan Wang, Xiaoning Li, Tien N. Nguyen, Shaohua Wang 等FSE 2024 · 被引用 6 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
