Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
Zhicheng Fang, Jingjie Zheng, Chenxu Fu, Wei Xu
Abstract
Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We introduce JAILBREAK FOUNDRY (JBF), a system that addresses this gap via a multi-agent workflow to translate jailbreak papers into executable modules for immediate evaluation within a unified harness. JBF features three core components: (i) JBF-LIB for shared contracts and reusable utilities; (ii) JBF-FORGE for the multi-agent paper-to-module translation; and (iii) JBF-EVAL for standardizing evaluations. Across 30 reproduced attacks, JBF achieves high fidelity with a mean (reproducedreported) attack success rate (ASR) deviation of percentage points. By leveraging shared infrastructure, JBF reduces attack-specific implementation code by nearly half relative to original repositories and achieves an 82.5% mean reused-code ratio. This system enables a standardized AdvBench evaluation of all 30 attacks across 10 victim models using a consistent GPT-4o judge. By automating both attack integration and standardized evaluation, JBF offers a scalable solution for creating living benchmarks that keep pace with the rapidly shifting security landscape.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 733f2e8b-8840-4d37-a37b-e9f8036c560aBuilds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 86 citations
Related papers
- From LLMs to MLLMs: Exploring the Landscape of Multimodal JailbreakingSiyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu WeiEMNLP 2024 · 5 citations
- GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak MethodsRuixuan Huang, Xunguang Wang, Zongjie Li, Daoyuan Wu et al.ICLR 2026 · 11 citations
- SoK: Robustness in Large Language Models against Jailbreak AttacksFeiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang et al.S&P 2026 · 4 citations
- LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesHaoyang Li, Huan Gao, Zhiyuan Zhao, Zhiyu Lin et al.ACL 2025
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack ExperienceXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li et al.EMNLP 2025 · 1 citation
