Agent Incentives: A Causal Perspective
Tom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega, Shane Legg
Abstract
We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a new graphical criterion for value of control, establishing its soundness and completeness. We also introduce two new concepts for incentive analysis: response incentives indicate which changes in the environment affect an optimal decision, while instrumental control incentives establish whether an agent can influence its utility via a variable X. For both new concepts, we provide sound and complete graphical criteria. We show by example how these results can help with evaluating the safety and fairness of an AI system. Race High school Education Grade Predicted grade Gender Accuracy
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d54dac6-79f6-42f2-bf7d-3946e99a2543Cited by top-tier papers13
- Robust agents learn causal world modelsJonathan Richens, Tom EverittICLR 2024 · 78 citations
- Honesty Is the Best Policy: Defining and Mitigating AI DeceptionFrancis Ward, Francesca Toni, Francesco Belardinelli, Tom EverittNeurIPS 2023 · 60 citations
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell et al.ICML 2024 · 44 citations
- Why Fair Labels Can Yield Unfair Predictions: Graphical Conditions for Introduced UnfairnessCarolyn Ashurst, Ryan Carey, Silvia Chiappa, Tom EverittAAAI 2022 · 17 citations
- Partial Counterfactual Identification of Continuous Outcomes with a Curvature Sensitivity ModelValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelNeurIPS 2023 · 15 citations
Builds on4
- A Calculus for Stochastic Interventions: Causal Effect Identification and Surrogate ExperimentsJuan D. Correa, Elias BareinboimAAAI 2020 · 90 citations
- Characterizing Optimal Mixed Policies: Where to Intervene and What to ObserveSanghack Lee, Elias BareinboimNeurIPS 2020 · 42 citations
- Asymptotically Unambitious Artificial General IntelligenceMichael K. Cohen, Badri N. Vellambi, Marcus HutterAAAI 2020 · 21 citations
- How RL Agents Behave When Their Actions Are ModifiedEric D. Langlois, Tom EverittAAAI 2021 · 17 citations
Related papers
- A Complete Criterion for Value of Information in Soluble Influence DiagramsChris van Merwijk, Ryan Carey, Tom EverittAAAI 2022 · 7 citations
- Path-Specific Objectives for Safer Agent IncentivesSebastian Farquhar, Ryan Carey, Tom EverittAAAI 2022 · 30 citations
- Causal Fairness for Outcome ControlDrago Plecko, Elias BareinboimNeurIPS 2023 · 17 citations
- Hedging as Reward Augmentation in Probabilistic Graphical ModelsDebarun Bhattacharjya, Radu MarinescuNeurIPS 2022
- Bayesian Active Learning for Bivariate Causal DiscoveryYuxuan Wang, Mingzhou Liu, Xinwei Sun, Wei Wang et al.ICML 2025
