Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied Agents
Seohui Bae, Jeonghye Kim, Youngchul Sung, Woohyung Lim
Abstract
In this paper, we propose a test-time adaptive agent that performs exploratory inference through posterior-guided belief refinement without relying on gradient-based updates or additional training for LLM agent operating under partial observability. Our agent maintains an external structured belief over the environment state, iteratively updates it via action-conditioned observations, and selects actions by maximizing predicted information gain over the belief space. We estimate information gain using a lightweight LLM-based surrogate and assess world alignment through a novel reward that quantifies the consistency between posterior belief and ground-truth environment configuration. Experiments show that our method outperforms inference-time scaling baselines such as prompt-augmented or retrieval-enhanced LLMs, in aligning with latent world states with significantly lower integration overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
Related papers
- Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy OptimizationXingyuan Hua, Sheng Yue, Ju RenICML 2026 · 1 citation
- AdaMEM: Test-Time Adaptive Memory for Language AgentsYunxiang Zhang, Yiheng Li, Ali Payani, Lu WangICML 2026
- Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information ForagingHongjin Qian, Zheng LiuNeurIPS 2025 · 24 citations
- Provable and Practical In-Context Policy Optimization for Self-ImprovementTianrun Yu, yuxiao Yang, Zhaoyang Wang, Kaixiang Zhao et al.ICLR 2026 · 1 citation
- Test-Time Adaptation for LLM Agents via Environment InteractionArthur Chen, Zuxin Liu, Jianguo Zhang, Akshara Prabhakar et al.ICLR 2026 · 18 citations
