Combining Cost Constrained Runtime Monitors for AI Safety
Tim Tian Hua, James Baskerville, Henri Lemoine, Mia Hopman, Aryan Bhatt, Tyler Tracy
Abstract
Monitoring AIs at runtime can help us detect and stop harmful actions. In this paper, we study how to efficiently combine multiple runtime monitors into a single monitoring protocol. The protocol's objective is to maximize the probability of applying a safety intervention on misaligned outputs (i.e., maximize recall). Since running monitors and applying safety interventions are costly, the protocol also needs to adhere to an average-case budget constraint. Taking the monitors' performance and cost as given, we develop an algorithm to find the best protocol. The algorithm exhaustively searches over when and which monitors to call, and allocates safety interventions based on the Neyman-Pearson lemma. By focusing on likelihood ratios and strategically trading off spending on monitors against spending on interventions, we more than double our recall rate compared to a naive baseline in a code review setting. We also show that combining two monitors can Pareto dominate using either monitor alone. Our framework provides a principled methodology for combining existing monitors to detect undesirable behavior in cost-sensitive settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes et al.NeurIPS 2025 · 52 citations
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal JailbreaksHoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic et al.ICLR 2026 · 39 citations
- Beyond Linear Probes: Dynamic Safety Monitoring for Language ModelsJames Oldfield, Philip Torr, Ioannis Patras, Adel Bibi et al.ICLR 2026 · 16 citations
Builds on5
- AI Control: Improving Safety Despite Intentional SubversionRyan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien RogerICML 2024 · 137 citations
- Detecting High-Stakes Interactions with Activation ProbesAlex McKenzie, Urja Pawar, Phil Blandfort, William Bankes et al.NeurIPS 2025 · 52 citations
- Control Tax: The Price of Keeping AI in CheckMikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel AlbanieICLR 2026 · 8 citations
- Adaptive Deployment of Untrusted LLMs Reduces Distributed ThreatsJiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt et al.ICLR 2025
- Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language ModelsGuobin Shen, Dongcheng Zhao, Yiting Dong, Xiang He et al.ICLR 2025
Related papers
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky et al.NeurIPS 2025 · 50 citations
- Quantitative and Approximate MonitoringThomas A. Henzinger, N. Ege SaraçLICS 2021 · 15 citations
- Delegated ClassificationEden Saig, Inbal Talgam-Cohen, Nir RosenfeldNeurIPS 2023 · 19 citations
- A Unifying Post-Processing Framework for Multi-Objective Learn-to-Defer ProblemsMohammad-Amin Charusaie, Samira SamadiNeurIPS 2024 · 6 citations
- Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and RiskTianrui Chen, Aditya Gangrade, Venkatesh SaligramaICML 2022 · 18 citations
