Estimating Worst-Case Frontier Risks of Open-Weight LLMs
Eric Wallace, Olivia Watkins, Miles Wang, Kai Chen, Chris Koch
摘要
In this paper, we study the worst-case frontier risks of the OpenAI gpt-oss model. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity. To maximize biological risk (biorisk), we curate tasks related to threat creation and train gpt-oss in an RL environment with web browsing. To maximize cybersecurity risk, we train gpt-oss in an agentic coding environment to solve capture-the-flag (CTF) challenges. We compare these MFT models against open- and closed-weight LLMs on frontier risk evaluations. Compared to frontier closed-weight models, MFT gpt-oss underperforms OpenAI o3, a model that is below Preparedness High capability level for biorisk and cybersecurity. Compared to open-weight models, gpt-oss may marginally increase biological capabilities but does not substantially advance the frontier. Taken together, these results led us to believe that the net new harm from releasing gpt-oss is limited, and we hope that our MFT approach can serve as useful guidance for estimating harm from future open-weight releases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning PerturbationYibo Wang, Tiansheng Huang, Li Shen, Huanjin Yao 等NeurIPS 2025 · 被引用 22 次
- Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient InfluenceQuoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui 等ICLR 2026 · 被引用 10 次
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-TuningWeitao Feng, Lixu Wang, Peizhuo Lv, Tianyi Wei 等CCS 2026 · 被引用 9 次
- Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention SinkGuozhi Liu, Weiwei Lin, Tiansheng Huang, Ruichao Mo 等ICML 2026 · 被引用 5 次
- Estimating Tail Risks in Language Model Output DistributionsRico Angell, Raghav Singhal, Zachary Horvitz, Zhou Yu 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper6
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue 等ICML 2024 · 被引用 390 次
- Covert Malicious Finetuning: Challenges in Safeguarding LLM AdaptationDanny Halawi, Alexander Wei, Eric Wallace, Tony Tong Wang 等ICML 2024 · 被引用 77 次
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak 等ICLR 2026 · 被引用 59 次
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li 等ICML 2025 · 被引用 1 次
- Adversaries Can Misuse Combinations of Safe ModelsErik Jones, Anca D. Dragan, Jacob SteinhardtICML 2025
相关 Paper
- ABC-Bench: An Agentic Bio-Capabilities Benchmark for BiosecurityAndrew Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman 等ICML 2026 · 被引用 6 次
- Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded OutputsJackson Kaunismaa, John Hughes, Christina Q. Knight, Avery Griffin 等ICLR 2026 · 被引用 7 次
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak SusceptibilityBrendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng 等EMNLP 2025 · 被引用 1 次
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety MechanismsJoshua Kazdan, Abhay Puri, Rylan Schaeffer, Lisa Yu 等ICLR 2026
