Agent Incentives: A Causal Perspective
Tom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega, Shane Legg
摘要
We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a new graphical criterion for value of control, establishing its soundness and completeness. We also introduce two new concepts for incentive analysis: response incentives indicate which changes in the environment affect an optimal decision, while instrumental control incentives establish whether an agent can influence its utility via a variable X. For both new concepts, we provide sound and complete graphical criteria. We show by example how these results can help with evaluating the safety and fairness of an AI system. Race High school Education Grade Predicted grade Gender Accuracy
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Robust agents learn causal world modelsJonathan Richens, Tom EverittICLR 2024 · 被引用 78 次
- Honesty Is the Best Policy: Defining and Mitigating AI DeceptionFrancis Ward, Francesca Toni, Francesco Belardinelli, Tom EverittNeurIPS 2023 · 被引用 60 次
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell 等ICML 2024 · 被引用 44 次
- Why Fair Labels Can Yield Unfair Predictions: Graphical Conditions for Introduced UnfairnessCarolyn Ashurst, Ryan Carey, Silvia Chiappa, Tom EverittAAAI 2022 · 被引用 17 次
- Partial Counterfactual Identification of Continuous Outcomes with a Curvature Sensitivity ModelValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelNeurIPS 2023 · 被引用 15 次
它引用的顶会 Paper4
- A Calculus for Stochastic Interventions: Causal Effect Identification and Surrogate ExperimentsJuan D. Correa, Elias BareinboimAAAI 2020 · 被引用 90 次
- Characterizing Optimal Mixed Policies: Where to Intervene and What to ObserveSanghack Lee, Elias BareinboimNeurIPS 2020 · 被引用 42 次
- Asymptotically Unambitious Artificial General IntelligenceMichael K. Cohen, Badri N. Vellambi, Marcus HutterAAAI 2020 · 被引用 21 次
- How RL Agents Behave When Their Actions Are ModifiedEric D. Langlois, Tom EverittAAAI 2021 · 被引用 17 次
相关 Paper
- A Complete Criterion for Value of Information in Soluble Influence DiagramsChris van Merwijk, Ryan Carey, Tom EverittAAAI 2022 · 被引用 7 次
- Path-Specific Objectives for Safer Agent IncentivesSebastian Farquhar, Ryan Carey, Tom EverittAAAI 2022 · 被引用 30 次
- Causal Fairness for Outcome ControlDrago Plecko, Elias BareinboimNeurIPS 2023 · 被引用 17 次
- Hedging as Reward Augmentation in Probabilistic Graphical ModelsDebarun Bhattacharjya, Radu MarinescuNeurIPS 2022
- Bayesian Active Learning for Bivariate Causal DiscoveryYuxuan Wang, Mingzhou Liu, Xinwei Sun, Wei Wang 等ICML 2025
