DOBF: A Deobfuscation Pre-Training Objective for Programming Languages
Marie-Anne Lachaux, Baptiste Rozière, Marc Szafraniec, Guillaume Lample
摘要
Recent advances in self-supervised learning have dramatically improved the state of the art on a wide variety of tasks. However, research in language model pretraining has mostly focused on natural languages, and it is unclear whether models like BERT and its variants provide the best pre-training when applied to other modalities, such as source code. In this paper, we introduce a new pre-training objective, DOBF, that leverages the structural aspect of programming languages and pre-trains a model to recover the original version of obfuscated source code. We show that models pre-trained with DOBF significantly outperform existing approaches on multiple downstream tasks, providing relative improvements of up to 12.2% in unsupervised code translation, and 5.3% in natural language code search. Incidentally, we found that our pre-trained model is able to deobfuscate fully obfuscated source files, and to suggest descriptive variable names. * Equal contribution. The order was determined randomly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman 等ICLR 2022 · 被引用 161 次
- InCoder: A Generative Model for Code Infilling and SynthesisDaniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang 等ICLR 2023 · 被引用 140 次
- NatGen: generative pre-training by "naturalizing" source codeSaikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T. Devanbu 等FSE 2022 · 被引用 101 次
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
它引用的顶会 Paper6
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 被引用 179 次
- Statistical Deobfuscation of Android ApplicationsBenjamin Bichsel, Veselin Raychev, Petar Tsankov, Martin T. VechevCCS 2016 · 被引用 128 次
- Neural reverse engineering of stripped binaries using augmented control flow graphsYaniv David, Uri Alon, Eran YahavOOPSLA 2020 · 被引用 83 次
相关 Paper
- ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation GroundingIndraneil Paul, Haoyi Yang, Goran Glavas, Kristian Kersting 等ICLR 2025
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Bridging Pre-trained Models and Downstream Tasks for Source Code UnderstandingDeze Wang, Zhouyang Jia, Shanshan Li, Yue Yu 等ICSE 2022 · 被引用 68 次
- Automating Code-Related Tasks Through Transformers: The Impact of Pre-trainingRosalia Tufano, Luca Pascarella, Gabriele BavotaICSE 2023 · 被引用 15 次
- ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningShangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng 等ICSE 2023 · 被引用 56 次
