MAGIKA: AI-Powered Content-Type Detection
Yanick Fratantonio, Luca Invernizzi, Loua Farah, Kurt Thomas, Marina Zhang, Ange Albertini, Francois Galilee, Giancarlo Metitieri, Julien Cretin, Alex Petit-Bianco, David Tao, Elie Bursztein
Abstract
The task of content-type detection-which entails identifying the data encoded in an arbitrary byte sequence-is critical for operating systems, development, reverse engineering environments, and a variety of security applications. In this paper, we introduce MAGIKA, a novel AI-powered content-type detection tool. Under the hood, MAGIKA employs a deep learning model that can execute on a single CPU with just 1MB of memory to store the model's weights. We show that MAGIKA achieves an average F1 score of 99% across over a hundred content types and a test set of more than 1M files, outperforming all existing content-type detection tools today. In order to foster adoption and improvements, we open source MAGIKA under an Apache 2 license on GitHub and will make our model and training pipeline publicly available. Our tool has already seen adoption by the Gmail email provider for attachment scanning, and it has been integrated with VirusTotal to aid with malware analysis. We note that this paper discusses the first iteration of MAGIKA, and a more recent version already supports more than 200 content types. The interested reader can see the latest development on the MAGIKA GitHub repository, available at github.com/google/magika.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c71e97a-d165-42b1-bb92-add4a90d4a67Related papers
- Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial AttacksMilad Nasr, Yanick Fratantonio, Luca Invernizzi, Ange Albertini et al.CCS 2025
- WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images DetectionYan Hong, Jianming Feng, Haoxing Chen, Jun Lan et al.AAAI 2025 · 13 citations
- Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment DetectorsJiahe Zhang, Jianjun Chen, Qi Wang, Hangyu Zhang et al.CCS 2024 · 2 citations
- MalwareTotal: Multi-Faceted and Sequence-Aware Bypass Tactics against Static Malware DetectionShuai He, Cai Fu, Hong Hu, Jiahe Chen et al.ICSE 2024 · 3 citations
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 4 citations
