ESCORT: Efficient Stein-variational and Sliced Consistency-Optimized Temporal Belief Representation for POMDPs
Yunuo Zhang, Baiting Luo, Ayan Mukhopadhyay, Gabor Karsai, Abhishek Dubey
Abstract
In Partially Observable Markov Decision Processes (POMDPs), maintaining and updating belief distributions over possible underlying states provides a principled way to summarize action-observation history for effective decision-making under uncertainty. As environments grow more realistic, belief distributions develop complexity that standard mathematical models cannot accurately capture, creating a fundamental challenge in maintaining representational accuracy. Despite advances in deep learning and probabilistic modeling, existing POMDP belief approximation methods fail to accurately represent complex uncertainty structures such as high-dimensional, multi-modal belief distributions, resulting in estimation errors that lead to suboptimal agent behaviors. To address this challenge, we present ESCORT (Efficient Stein-variational and sliced Consistency-Optimized Representation for Temporal beliefs), a particle-based framework for capturing complex, multi-modal distributions in high-dimensional belief spaces. ESCORT extends SVGD with two key innovations: correlation-aware projections that model dependencies between state dimensions, and temporal consistency constraints that stabilize updates while preserving correlation structures. This approach retains SVGD's attractive-repulsive particle dynamics while enabling accurate modeling of intricate correlation patterns. Unlike particle filters prone to degeneracy or parametric methods with fixed representational capacity, ESCORT dynamically adapts to belief landscape complexity without resampling or restrictive distributional assumptions. We demonstrate ESCORT's effectiveness through extensive evaluations on both POMDP domains and synthetic multi-modal distributions of varying dimensionality, where it consistently outperforms state-of-theart methods in terms of belief approximation accuracy and downstream decision quality. Our code is available at https://github.com/scope-lab-vu/ESCORT.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Projected Stein Variational Gradient DescentPeng Chen, Omar GhattasNeurIPS 2020 · 84 citations
- Discriminative Particle Filter Reinforcement Learning for Complex Partial observationsXiao Ma, Péter Karkus, David Hsu, Wee Sun Lee et al.ICLR 2020 · 50 citations
- Adaptive Online Packing-guided Search for POMDPsChenyang Wu, Guoyu Yang, Zongzhang Zhang, Yang Yu et al.NeurIPS 2021 · 28 citations
- Flow-based Recurrent Belief State Learning for POMDPsXiaoyu Chen, Yao Mark Mu, Ping Luo, Shengbo Li et al.ICML 2022 · 26 citations
- Set-membership Belief State-based Reinforcement Learning for POMDPsWei Wei, Lijun Zhang, Lin Li, Huizhong Song et al.ICML 2023 · 2 citations
Related papers
- Search and Explore: Symbiotic Policy Synthesis in POMDPsRoman Andriushchenko, Alexander Bork, Milan Ceska, Sebastian Junges et al.CAV 2023 · 7 citations
- The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space ModelsRaphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez et al.ICLR 2024 · 10 citations
- Factored Online Planning in Many-Agent POMDPsMaris F. L. Galesloot, Thiago D. Simão, Sebastian Junges, Nils JansenAAAI 2024 · 3 citations
- Provably Efficient Representation Learning with Tractable Planning in Low-Rank POMDPJiacheng Guo, Zihao Li, Huazheng Wang, Mengdi Wang et al.ICML 2023 · 8 citations
- Sequential Monte Carlo for Policy Optimization in Continuous POMDPsHany Abdulsamad, Sahel Mohammad Iqbal, Simo SärkkäNeurIPS 2025 · 4 citations
