ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs
Yunuo Zhang, Baiting Luo, Ayan Mukhopadhyay, Gabor Karsai, Abhishek Dubey
摘要
In Partially Observable Markov Decision Processes (POMDPs), maintaining and updating belief distributions over possible underlying states provides a principled way to summarize action-observation history for effective decision-making under uncertainty. As environments grow more realistic, belief distributions develop complexity that standard mathematical models cannot accurately capture, creating a fundamental challenge in maintaining representational accuracy. Despite advances in deep learning and probabilistic modeling, existing POMDP belief approximation methods fail to accurately represent complex uncertainty structures such as high-dimensional, multi-modal belief distributions, resulting in estimation errors that lead to suboptimal agent behaviors. To address this challenge, we present ESCORT (Efficient Stein-variational and sliced Consistency-Optimized Representation for Temporal beliefs), a particle-based framework for capturing complex, multi-modal distributions in high-dimensional belief spaces. ESCORT extends SVGD with two key innovations: correlation-aware projections that model dependencies between state dimensions, and temporal consistency constraints that stabilize updates while preserving correlation structures. This approach retains SVGD's attractive-repulsive particle dynamics while enabling accurate modeling of intricate correlation patterns. Unlike particle filters prone to degeneracy or parametric methods with fixed representational capacity, ESCORT dynamically adapts to belief landscape complexity without resampling or restrictive distributional assumptions. We demonstrate ESCORT's effectiveness through extensive evaluations on both POMDP domains and synthetic multi-modal distributions of varying dimensionality, where it consistently outperforms state-of-theart methods in terms of belief approximation accuracy and downstream decision quality. Our code is available at https://github.com/scope-lab-vu/ESCORT.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Projected Stein Variational Gradient DescentPeng Chen, Omar GhattasNeurIPS 2020 · 被引用 84 次
- Discriminative Particle Filter Reinforcement Learning for Complex Partial observationsXiao Ma, Péter Karkus, David Hsu, Wee Sun Lee 等ICLR 2020 · 被引用 50 次
- Adaptive Online Packing-guided Search for POMDPsChenyang Wu, Guoyu Yang, Zongzhang Zhang, Yang Yu 等NeurIPS 2021 · 被引用 28 次
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li 等ICML 2022 · 被引用 26 次
- Set-membership Belief State-based Reinforcement Learning for POMDPsWei Wei, Lijun Zhang, Lin Li, Huizhong Song 等ICML 2023 · 被引用 2 次
相关 Paper
- Search and Explore: Symbiotic Policy Synthesis in POMDPsRoman Andriushchenko, Alexander Bork, Milan Ceska, Sebastian Junges 等CAV 2023 · 被引用 7 次
- The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space ModelsRaphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez 等ICLR 2024 · 被引用 10 次
- Factored Online Planning in Many-Agent POMDPsMaris F. L. Galesloot, Thiago D. Simão, Sebastian Junges, Nils JansenAAAI 2024 · 被引用 3 次
- Provably Efficient Representation Learning with Tractable Planning in Low-Rank POMDPJiacheng Guo, Zihao Li, Huazheng Wang, Mengdi Wang 等ICML 2023 · 被引用 8 次
- Sequential Monte Carlo for Policy Optimization in Continuous POMDPsHany Abdulsamad, Sahel Mohammad Iqbal, Simo SärkkäNeurIPS 2025 · 被引用 4 次
