Scaling Laws For Scalable Oversight
Joshua Engels, David D. Baek, Subhash Kantamneni, Max Tegmark
摘要
Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is still unclear how scalable oversight itself scales. To address this gap, we propose a framework that quantifies the probability of successful oversight as a function of the capabilities of the overseer and the system being overseen. Specifically, our framework models oversight as a game between capability-mismatched players; the players have oversight-specific Elo scores that are a piecewise-linear function of their general intelligence, with two plateaus corresponding to task incompetence and task saturation. We validate our framework with a modified version of the game Nim and then apply it to four oversight games: Mafia, Debate, Backdoor Code and Wargames. For each game, we find scaling laws that approximate how domain performance depends on general AI system capability. We then build on our findings in a theoretical study of Nested Scalable Oversight (NSO), a process in which trusted models oversee untrusted stronger models, which then become the trusted models in the next step. We identify conditions under which NSO succeeds and derive numerically (and in some cases analytically) the optimal number of oversight levels to maximize the probability of oversight success. We also apply our theory to our four oversight games, where we find that NSO success rates at a general Elo gap of 400 are 13.5% for Mafia, 51.7% for Debate, 10.0% for Backdoor Code, and 9.4% for Wargames; these rates decline further when overseeing stronger systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre 等ICLR 2026 · 被引用 26 次
- Control Tax: The Price of Keeping AI in CheckMikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel AlbanieICLR 2026 · 被引用 8 次
- Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge RegressionDiyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco MondelliICML 2026
- Calibrating Conservatism for Scalable OversightWilliam Overman, Mohsen BayatiICML 2026
它引用的顶会 Paper11
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- Debating with More Persuasive LLMs Leads to More Truthful AnswersAkbir Khan, John Hughes, Dan Valentine, Laura Ruis 等ICML 2024 · 被引用 244 次
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
- Scaling Laws for Generative Mixed-Modal Language ModelsArmen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu 等ICML 2023 · 被引用 149 次
相关 Paper
- On scalable oversight with weak LLMs judging strong LLMsZachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen 等NeurIPS 2024 · 被引用 116 次
- Great Models Think Alike and this Undermines AI OversightShashwat Goel, Joschka Strüber, Ilze Amanda Auzina, Karuna K. Chandra 等ICML 2025
- A theoretical case-study of Scalable Oversight in Hierarchical Reinforcement LearningTom Yan, Zachary C. LiptonNeurIPS 2024 · 被引用 3 次
- Towards Scalable Oversight via Partitioned Human SupervisionRen Yin, Takashi Ishida, Masashi SugiyamaICLR 2026
- Governing AI Agent with Capability and Delegation ChainJiaqin Yan, Guoxing Chen, Yan Meng, Haojin ZhuCCS 2026
