MAGIKA: AI-Powered Content-Type Detection
Yanick Fratantonio, Luca Invernizzi, Loua Farah, Kurt Thomas, Marina Zhang, Ange Albertini, Francois Galilee, Giancarlo Metitieri, Julien Cretin, Alex Petit-Bianco, David Tao, Elie Bursztein
摘要
The task of content-type detection-which entails identifying the data encoded in an arbitrary byte sequence-is critical for operating systems, development, reverse engineering environments, and a variety of security applications. In this paper, we introduce MAGIKA, a novel AI-powered content-type detection tool. Under the hood, MAGIKA employs a deep learning model that can execute on a single CPU with just 1MB of memory to store the model's weights. We show that MAGIKA achieves an average F1 score of 99% across over a hundred content types and a test set of more than 1M files, outperforming all existing content-type detection tools today. In order to foster adoption and improvements, we open source MAGIKA under an Apache 2 license on GitHub and will make our model and training pipeline publicly available. Our tool has already seen adoption by the Gmail email provider for attachment scanning, and it has been integrated with VirusTotal to aid with malware analysis. We note that this paper discusses the first iteration of MAGIKA, and a more recent version already supports more than 200 content types. The interested reader can see the latest development on the MAGIKA GitHub repository, available at github.com/google/magika.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial AttacksMilad Nasr, Yanick Fratantonio, Luca Invernizzi, Ange Albertini 等CCS 2025
- WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images DetectionYan Hong, Jianming Feng, Haoxing Chen, Jun Lan 等AAAI 2025 · 被引用 13 次
- Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment DetectorsJiahe Zhang, Jianjun Chen, Qi Wang, Hangyu Zhang 等CCS 2024 · 被引用 2 次
- MalwareTotal: Multi-Faceted and Sequence-Aware Bypass Tactics against Static Malware DetectionShuai He, Cai Fu, Hong Hu, Jiahe Chen 等ICSE 2024 · 被引用 3 次
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 被引用 4 次
