Dynamic Knowledge Injection for AIXI Agents
Samuel Yang-Zhao, Kee Siong Ng, Marcus Hutter
摘要
Prior approximations of AIXI, a Bayesian optimality notion for general reinforcement learning, can only approximate AIXI's Bayesian environment model using an a-priori defined set of models. This is a fundamental source of epistemic uncertainty for the agent in settings where the existence of systematic bias in the predefined model class cannot be resolved by simply collecting more data from the environment. We address this issue in the context of Human-AI teaming by considering a setup where additional knowledge for the agent in the form of new candidate models arrives from a human operator in an online fashion. We introduce a new agent called DynamicHedgeAIXI that maintains an exact Bayesian mixture over dynamically changing sets of models via a time-adaptive prior constructed from a variant of the Hedge algorithm. The DynamicHedgeAIXI agent is the richest direct approximation of AIXI known to date and comes with good performance guarantees. Experimental results on epidemic control on contact networks validates the agent's practical utility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Hedging as Reward Augmentation in Probabilistic Graphical ModelsDebarun Bhattacharjya, Radu MarinescuNeurIPS 2022
- Epistemic Gain, Aleatoric Cost: Uncertainty Decomposition in Multi-Agent Debate for Math ReasoningDan Qiao, Binbin Chen, Fengyu Cai, Jianlong Chen 等ICML 2026 · 被引用 3 次
- Regularized Offline Policy Optimization with Posterior Hybrid Bayesian BeliefHongqiang Lin, Pengfei Wang, Nenggan ZhengICML 2026 · 被引用 1 次
- Explicable Policy SearchZe Gong, Yu ZhangNeurIPS 2022 · 被引用 5 次
- Collective Intelligence in Human-AI Teams: A Bayesian Theory of Mind ApproachSamuel Westby, Christoph RiedlAAAI 2023 · 被引用 35 次
