Calibrating Conservatism for Scalable Oversight
William Overman, Mohsen Bayati
Abstract
Agentic AI systems capable of autonomous planning and extended environmental interaction pose a fundamental control problem: how can humans maintain meaningful oversight of systems that may exceed human capabilities? While scalable oversight is widely studied, existing approaches often rely on complex assumptions, remain largely heuristic, or lack practical methods for sequential settings with statistical guarantees. We introduce Calibrated Collective Oversight (CCO), which aggregates diverse auxiliary scoring functions into a penalty that measures deviation from a conservative baseline. Inspired by Attainable Utility Preservation, CCO enables collective conservatism: when multiple oversight signals register concern, the agent defers. CCO calibrates this conservatism online using Conformal Decision Theory, ensuring that undesirable outcomes remain below a user-specified target with finite-time bounds and no distributional assumptions. Experiments on SWE-bench demonstrate that weaker overseers successfully constrain an adversarially misaligned stronger agent. Similarly, on MACHIAVELLI, CCO achieves substantial reductions in ethical violations while preserving reward. In both settings, empirical violation rates closely match the specified targets. Our work demonstrates that combining penalty-based conservatism with online calibration yields practical oversight with statistical guarantees suited for agentic deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf69bc4c-abf7-4eeb-960a-ddbe7562a2b2Builds on16
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- Conformal Risk ControlAnastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei et al.ICLR 2024 · 242 citations
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkAlexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li et al.ICML 2023 · 200 citations
Related papers
- Conformal Policy ControlDrew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho et al.ICML 2026 · 3 citations
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language ModelsWilliam Overman, Mohsen BayatiNeurIPS 2025 · 12 citations
- Scaling Laws For Scalable OversightJoshua Engels, David D. Baek, Subhash Kantamneni, Max TegmarkNeurIPS 2025 · 22 citations
- The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and AutonomyWilliam Overman, Mohsen BayatiICML 2026 · 7 citations
- MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward HackingSebastian Farquhar, Vikrant Varma, David Lindner, David K. Elson et al.ICML 2025
