Asymptotically Unambitious Artificial General Intelligence
Michael K. Cohen, Badri N. Vellambi, Marcus Hutter
摘要
General intelligence, the ability to solve arbitrary solvable problems, is supposed by many to be artificially constructible. Narrow intelligence, the ability to solve a given particularly difficult problem, has seen impressive recent development. Notable examples include self-driving cars, Go engines, image classifiers, and translators. Artificial General Intelligence (AGI) presents dangers that narrow intelligence does not: if something smarter than us across every domain were indifferent to our concerns, it would be an existential threat to humanity, just as we threaten many species despite no ill will. Even the theory of how to maintain the alignment of an AGI's goals with our own has proven highly elusive. We present the first algorithm we are aware of for asymptotically unambitious AGI, where “unambitiousness” includes not seeking arbitrary power. Thus, we identify an exception to the Instrumental Convergence Thesis, which is roughly that by default, an AGI would seek power, including over us.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Agent Incentives: A Causal PerspectiveTom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega 等AAAI 2021 · 被引用 66 次
- Logical Phase Transitions: Understanding Collapse in LLM Logical ReasoningXinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu 等ACL 2026 · 被引用 28 次
- Semantic-Aware Logical Reasoning via a Semiotic FrameworkYunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li 等ACL 2026 · 被引用 27 次
- A Complete Criterion for Value of Information in Soluble Influence DiagramsChris van Merwijk, Ryan Carey, Tom EverittAAAI 2022 · 被引用 7 次
相关 Paper
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIsMantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim 等NeurIPS 2025 · 被引用 84 次
- Advantage Alignment AlgorithmsJuan Agustin Duque, Milad Aghajohari, Tim Cooijmans, Razvan Ciuca 等ICLR 2025
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural NetworksSu Hyeong Lee, Risi Kondor, Richard NgoICML 2026
- On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI AlignmentSarah Ball, Greg Gluch, Shafi Goldwasser, Frauke Kreuter 等ICLR 2026 · 被引用 16 次
