Combining Cost Constrained Runtime Monitors for AI Safety
Tim Tian Hua, James Baskerville, Henri Lemoine, Mia Hopman, Aryan Bhatt, Tyler Tracy
摘要
Monitoring AIs at runtime can help us detect and stop harmful actions. In this paper, we study how to efficiently combine multiple runtime monitors into a single monitoring protocol. The protocol's objective is to maximize the probability of applying a safety intervention on misaligned outputs (i.e., maximize recall). Since running monitors and applying safety interventions are costly, the protocol also needs to adhere to an average-case budget constraint. Taking the monitors' performance and cost as given, we develop an algorithm to find the best protocol. The algorithm exhaustively searches over when and which monitors to call, and allocates safety interventions based on the Neyman-Pearson lemma. By focusing on likelihood ratios and strategically trading off spending on monitors against spending on interventions, we more than double our recall rate compared to a naive baseline in a code review setting. We also show that combining two monitors can Pareto dominate using either monitor alone. Our framework provides a principled methodology for combining existing monitors to detect undesirable behavior in cost-sensitive settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes 等NeurIPS 2025 · 被引用 52 次
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal JailbreaksHoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic 等ICLR 2026 · 被引用 39 次
- Beyond Linear Probes: Dynamic Safety Monitoring for Language ModelsJames Oldfield, Philip Torr, Ioannis Patras, Adel Bibi 等ICLR 2026 · 被引用 16 次
它引用的顶会 Paper5
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 被引用 137 次
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes 等NeurIPS 2025 · 被引用 52 次
- Control Tax: The Price of Keeping AI in CheckMikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel AlbanieICLR 2026 · 被引用 8 次
- Adaptive Deployment of Untrusted LLMs Reduces Distributed ThreatsJiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt 等ICLR 2025
- Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language ModelsGuobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He 等ICLR 2025
相关 Paper
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky 等NeurIPS 2025 · 被引用 50 次
- Quantitative and Approximate MonitoringThomas A. Henzinger, N. Ege SaraçLICS 2021 · 被引用 15 次
- Delegated ClassificationEden Saig, Inbal Talgam-Cohen, Nir RosenfeldNeurIPS 2023 · 被引用 19 次
- A Unifying Post-Processing Framework for Multi-Objective Learn-to-Defer ProblemsMohammad-Amin Charusaie, Samira SamadiNeurIPS 2024 · 被引用 6 次
- Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and RiskTianrui Chen, Aditya Gangrade, Venkatesh SaligramaICML 2022 · 被引用 18 次
