Estimating Worst-Case Frontier Risks of Open-Weight LLMs
Eric Wallace, Olivia Watkins, Miles Wang, Kai Chen, Chris Koch
Abstract
In this paper, we study the worst-case frontier risks of the OpenAI gpt-oss model. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity. To maximize biological risk (biorisk), we curate tasks related to threat creation and train gpt-oss in an RL environment with web browsing. To maximize cybersecurity risk, we train gpt-oss in an agentic coding environment to solve capture-the-flag (CTF) challenges. We compare these MFT models against open- and closed-weight LLMs on frontier risk evaluations. Compared to frontier closed-weight models, MFT gpt-oss underperforms OpenAI o3, a model that is below Preparedness High capability level for biorisk and cybersecurity. Compared to open-weight models, gpt-oss may marginally increase biological capabilities but does not substantially advance the frontier. Taken together, these results led us to believe that the net new harm from releasing gpt-oss is limited, and we hope that our MFT approach can serve as useful guidance for estimating harm from future open-weight releases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 571d499c-07d7-4eb6-8f5c-603b9a525cefCited by top-tier papers10
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning PerturbationYibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao et al.NeurIPS 2025 · 22 citations
- Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient InfluenceQuoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui et al.ICLR 2026 · 10 citations
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-TuningWeitao Feng, Lixu Wang, Peizhuo Lv, Tianyi Wei et al.CCS 2026 · 9 citations
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo et al.ICML 2026 · 5 citations
- Estimating Tail Risks in Language Model Output DistributionsRico Angell, Raghav Singhal, Zachary Horvitz, Zhou Yu et al.ICML 2026 · 3 citations
Builds on6
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Covert Malicious Finetuning: Challenges in Safeguarding LLM AdaptationDanny Halawi, Alexander Wei, Eric Wallace, Tony Tong Wang et al.ICML 2024 · 77 citations
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak et al.ICLR 2026 · 59 citations
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li et al.ICML 2025 · 1 citation
- Adversaries Can Misuse Combinations of Safe ModelsErik Jones, Anca D. Dragan, Jacob SteinhardtICML 2025
Related papers
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for BiosecurityAndrew Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman et al.ICML 2026 · 6 citations
- Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded OutputsJackson Kaunismaa, John Hughes, Christina Q. Knight, Avery Griffin et al.ICLR 2026 · 7 citations
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak SusceptibilityBrendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng et al.EMNLP 2025 · 1 citation
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety MechanismsJoshua Kazdan, Abhay Puri, Rylan Schaeffer, Lisa Yu et al.ICLR 2026
