A Fine-Tuning Approach to Belief State Modeling
Samuel Sokota, Hengyuan Hu, David J. Wu, J. Zico Kolter, Jakob Nicolaus Foerster, Noam Brown
摘要
We investigate the challenge of modeling the belief state of a partially observable Markov system, given sample-access to its dynamics model. This problem setting is often approached using parametric sequential generative modeling methods. However, these methods do not leverage any additional computation at inference time to increase their accuracy. Moreover, applying these methods to belief state modeling in certain multi-agent settings would require passing policies into the belief model-at the time of writing, there have been no successful demonstrations of this. Toward addressing these shortcomings, we propose an inference-time improvement framework for parametric sequential generative modeling methods called belief fine-tuning (BFT). BFT leverages approximate dynamic programming in the form of fine-tuning to determine the model parameters at each time step. It can improve the accuracy of the belief model at test time because it specializes the model to the space of local observations. Furthermore, because this specialization occurs after the action or policy has already been decided, BFT does not require the belief model to process it as input. As a result of the latter point, BFT enables, for the first time, approximate public belief state search in imperfect-information games where the number of possible information states is too large to track tabularly. We exhibit these findings on large-scale variants of the benchmark game Hanabi.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Permutation Equivariant Neural FunctionalsAllan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace 等NeurIPS 2023 · 被引用 84 次
- Monomial Matrix Group Equivariant Neural Functional NetworksHoang V. Tran, Thieu N. Vo, Tho Huu, An Nguyen The 等NeurIPS 2024 · 被引用 19 次
- Abstracting Imperfect Information Away from Two-Player Zero-Sum GamesSamuel Sokota, Ryan D'Orazio, Chun Kai Ling, David J. Wu 等ICML 2023 · 被引用 8 次
- The Update-Equivalence Framework for Decision-Time PlanningSamuel Sokota, Gabriele Farina, David J. Wu, Hengyuan Hu 等ICLR 2024 · 被引用 5 次
- Off-Team LearningBrandon Cui, Hengyuan Hu, Andrei Lupu, Samuel Sokota 等NeurIPS 2022 · 被引用 4 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 被引用 205 次
- Simplified Action Decoder for Deep Multi-Agent Reinforcement LearningHengyuan Hu, Jakob N. FoersterICLR 2020 · 被引用 88 次
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 被引用 87 次
相关 Paper
- Scalable Online Planning via Reinforcement Learning Fine-TuningArnaud Fickinger, Hengyuan Hu, Brandon Amos, Stuart J. Russell 等NeurIPS 2021 · 被引用 26 次
- Information Particle Filter Tree: An Online Algorithm for POMDPs with Belief-Based Rewards on Continuous DomainsJohannes Fischer, Ömer Sahin TasICML 2020 · 被引用 42 次
- Adaptive Online Packing-guided Search for POMDPsChenyang Wu, Guoyu Yang, Zongzhang Zhang, Yang Yu 等NeurIPS 2021 · 被引用 28 次
- Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied AgentsSeohui Bae, Jeonghye Kim, Youngchul Sung, Woohyung LimCVPR 2026 · 被引用 1 次
- GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent SystemYiqin Yang, Xu Yang, Yuhua Jiang, Ni Mu 等ICLR 2026
