MalCL: Leveraging GAN-Based Generative Replay to Combat Catastrophic Forgetting in Malware Classification
Jimin Park, AHyun Ji, Minji Park, Mohammad Saidur Rahman, Se Eun Oh
Abstract
Continual Learning (CL) for malware classification tackles the rapidly evolving nature of malware threats and the frequent emergence of new types. Generative Replay (GR)based CL systems utilize a generative model to produce synthetic versions of past data, which are then combined with new data to retrain the primary model. Traditional machine learning techniques in this domain often struggle with catastrophic forgetting, where a model's performance on old data degrades over time. In this paper, we introduce a GR-based CL system that employs Generative Adversarial Networks (GANs) with feature matching loss to generate high-quality malware samples. Additionally, we implement innovative selection schemes for replay samples based on the model's hidden representations. Our comprehensive evaluation across Windows and Android malware datasets in a class-incremental learning scenariowhere new classes are introduced continuously over multiple tasks -demonstrates substantial performance improvements over previous methods. For example, our system achieves an average accuracy of 55% on Windows malware samples, significantly outperforming other GR-based models by 28%. This study provides practical insights for advancing GR-based malware classification systems. The implementation is available at https://github.com/MalwareReplayGAN/ MalCL 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76ee1844-2e4a-49dc-a760-0aac827d8847Cited by top-tier papers2
- LAMDA: A Longitudinal Android Malware Benchmark for Concept Drift AnalysisMd Ahsanul Haque, Ismail Hossain, Md Mahmuduzzaman Kamol, Md Jahangir Alam et al.ICLR 2026 · 14 citations
- HyperGLLM: An Efficient Framework for Endpoint Threat Detection via Hypergraph-Enhanced Large Language ModelsHongyi Zhou, Jianfeng Pan, Min Peng, Shaomang Huang et al.AAAI 2026
Builds on6
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- DDGR: Continual Learning with Deep Diffusion-based Generative ReplayRui Gao, Weiwei LiuICML 2023 · 101 citations
- Classifying Sequences of Extreme Length with Constant Memory Applied to Malware DetectionEdward Raff, William Fleshman, Richard Zak, Hyrum S. Anderson et al.AAAI 2021 · 70 citations
- Augmented Memory Replay-based Continual Learning Approaches for Network Intrusion DetectionSuresh Kumar Amalapuram, Sumohana S. Channappayya, Bheemarjuna Reddy TammaNeurIPS 2023 · 42 citations
- Continuous Learning for Android Malware DetectionYizheng Chen, Zhoujie Ding, David A. WagnerUSENIX Security 2023
Related papers
- Guided Retraining to Enhance the Detection of Difficult Android MalwareNadia Daoudi, Kevin Allix, Tegawendé F. Bissyandé, Jacques KleinISSTA 2023 · 4 citations
- Semantics-Driven Generative Replay for Few-Shot Class Incremental LearningAishwarya Agarwal, Biplab Banerjee, Fabio Cuzzolin, Subhasis ChaudhuriACM MM 2022 · 26 citations
- GCR: Gradient Coreset based Replay Buffer Selection for Continual LearningRishabh Tiwari, KrishnaTeja Killamsetty, Rishabh K. Iyer, Pradeep ShenoyCVPR 2022 · 102 citations
- RECALL: Replay-based Continual Learning in Semantic SegmentationAndrea Maracani, Umberto Michieli, Marco Toldo, Pietro ZanuttighICCV 2021 · 148 citations
- DeepCollaboration: Collaborative Generative and Discriminative Models for Class Incremental LearningBo Cui, Guyue Hu, Shan YuAAAI 2021 · 11 citations
