Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
Su Hyeong Lee, Risi Kondor, Richard Ngo
Abstract
We develop a theory of intelligent agency grounded in probabilistic modeling for neural models. Agents are represented as outcome distributions with epistemic utility given by log score, and compositions are defined through weighted logarithmic pooling that strictly improves every member's welfare. We prove that strict unanimity is impossible under linear pooling or in binary outcome spaces, but possible with three or more outcomes. Our framework admits recursive structure via cloning invariance, continuity, and openness, while tilt-based analysis rules out trivial duplication. Finally, we formalize an agentic alignment phenomenon in LLMs using our theory: eliciting a benevolent persona ("Luigi'') induces an antagonistic counterpart ("Waluigi''), while a manifest-then-suppress Waluigi strategy yields strictly larger first-order misalignment reduction than pure Luigi reinforcement alone. These results clarify how developing a principled mathematical framework for how subagents can coalesce into coherent higher-level entities provides novel implications for alignment in agentic AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfabb257-079b-4422-8ee5-549e242d5024Builds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureGuangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang et al.ICLR 2024 · 311 citations
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
Related papers
- Emergent Coordination in Multi-Agent Language ModelsChristoph RiedlICLR 2026 · 25 citations
- Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM AgentsShuhui Zhu, Yue Lin, Shriya Kaistha, Wenhao Li et al.ICML 2026
- Emergent Alignment via CompetitionNatalie Collina, Surbhi Goel, Aaron Roth, Emily Ryu et al.ICML 2026
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIsMantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim et al.NeurIPS 2025 · 84 citations
- Advantage Alignment AlgorithmsJuan Agustin Duque, Milad Aghajohari, Tim Cooijmans, Razvan Ciuca et al.ICLR 2025
