Threat2Traffic: Multi-Agent Environment Synthesis for Malware Traffic Generation from Threat Intelligence
Haoyang Chen, Chang Liu, Zhong Guan, Junzheng Shi, Gaopeng Gou, Gang Xiong
Abstract
Data-driven cybersecurity research is fundamentally constrained by the scarcity of labeled datasets, yet acquiring authentic, large-scale malware traffic remains bottlenecked by obsolescent public datasets, unscalable manual construction, and inflexible sandboxes that fail to satisfy the sample-specific dependencies required for malware to exhibit malicious behavior. Threat intelligence documents these dependencies, and LLM agents offer a path to extract them for environment construction, yet directly applying such agents faces two challenges: input-side ambiguity and output-side fragility. In this paper, we propose Threat2Traffic, a multi-agent framework that extracts sample-specific dependencies from threat intelligence, reconstructs tailored environments, and captures malware traffic. To address input-side ambiguity, it formulates dependency extraction as structured multi-agent deliberation over an evidence graph. To overcome output-side fragility, it incorporates invariant-guided synthesis with dual-layer validation under syntactic and semantic constraints. Evaluated on 1,077 samples across eight malware families, Threat2Traffic achieves 83.1% reproduction success, highlighting its effectiveness for scalable and realistic malware traffic generation. We release the core source code and traffic dataset at https://github.com/apos3637/Threat2Traffic
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cfdf2b2-f7b2-4abf-9cc3-9c06c760c9aeBuilds on14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang et al.EMNLP 2024 · 177 citations
- Spotless Sandboxes: Evading Malware Analysis Systems Using Wear-and-Tear ArtifactsNajmeh Miramirkhani, Mahathi Priya Appini, Nick Nikiforakis, Michalis PolychronakisS&P 2017 · 134 citations
Related papers
- CVE-Genie: An LLM-Based Multi-Agent Framework for Reproducing CVEsSaad Ullah, Praneeth Balasubramanian, Wenbo Guo, Amanda Burnett et al.CCS 2026
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM AgentsPengfei He, Ash Fox, Lesly Miculicich, Stefan Friedli et al.ICML 2026 · 10 citations
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat InvestigationYiran Wu, Mauricio Velazco, Andrew Zhao, Manuel Luján et al.ICML 2026 · 14 citations
- AdvTG: An Adversarial Traffic Generation Framework to Deceive DL-Based Malicious Traffic Detection ModelsPeishuai Sun, Xiaochun Yun, Shuhao Li, Tao Yin et al.WWW 2025 · 5 citations
- GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph ModelingJialong Zhou, Lichao Wang, Xiao YangNeurIPS 2025 · 40 citations
