Robin: A Novel Method to Produce Robust Interpreters for Deep Learning-Based Code Classifiers
Zhen Li, Ruqian Zhang, Deqing Zou, Ning Wang, Yating Li, Shouhuai Xu, Chen Chen, Hai Jin
Abstract
Deep learning has been widely used in source code classification tasks, such as code classification according to their functionalities, code authorship attribution, and vulnerability detection. Unfortunately, the black-box nature of deep learning makes it hard to interpret and understand why a classifier (i.e., classification model) makes a particular prediction on a given example. This lack of interpretability (or explainability) might have hindered their adoption by practitioners because it is not clear when they should or should not trust a classifier's prediction. The lack of interpretability has motivated a number of studies in recent years. However, existing methods are neither robust nor able to cope with out-of-distribution examples. In this paper, we propose a novel method to produce Robust interpreters for a given deep learning-based code classifier; the method is dubbed Robin. The key idea behind Robin is a novel hybrid structure combining an interpreter and two approximators, while leveraging the ideas of adversarial training and data augmentation. Experimental results show that on average the interpreter produced by Robin achieves a 6.11% higher fidelity (evaluated on the classifier), 67.22% higher fidelity (evaluated on the approximator), and 15.87x higher robustness than that of the three existing interpreters we evaluated. Moreover, the interpreter is 47.31% less affected by out-of-distribution examples than that of LEMNA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5018f44c-8a8c-4811-92e5-9804f84c74ebCited by top-tier papers2
- Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and MemorizationZhi Chen, Lingxiao JiangASE 2024 · 4 citations
- Snopy: Bridging Sample Denoising with Causal Graph Learning for Effective Vulnerability DetectionSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo et al.ASE 2024 · 2 citations
Builds on15
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su et al.CCS 2018 · 336 citations
- Robust Counterfactual Explanations on Graph Neural NetworksMohit Bajaj, Lingyang Chu, Zi Yu Xue, Jian Pei et al.NeurIPS 2021 · 140 citations
- Misleading Authorship Attribution of Source Code using Adversarial LearningErwin Quiring, Alwin Maier, Konrad RieckUSENIX Security 2019 · 123 citations
- Large-Scale and Language-Oblivious Code Authorship IdentificationMohammed Abuhamad, Tamer AbuHmed, Aziz Mohaisen, DaeHun NyangCCS 2018 · 102 citations
Related papers
- One step further: evaluating interpreters using metamorphic testingMing Fan, Jiali Wei, Wuxia Jin, Zhou Xu et al.ISSTA 2022 · 7 citations
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 27 citations
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
- Enhancing Robustness of Code Authorship Attribution through Expert Feature KnowledgeXiaowei Guo, Cai Fu, Juan Chen, Hongle Liu et al.ISSTA 2024 · 2 citations
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 107 citations
