Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied Agents
Seohui Bae, Jeonghye Kim, Youngchul Sung, Woohyung Lim
摘要
In this paper, we propose a test-time adaptive agent that performs exploratory inference through posterior-guided belief refinement without relying on gradient-based updates or additional training for LLM agent operating under partial observability. Our agent maintains an external structured belief over the environment state, iteratively updates it via action-conditioned observations, and selects actions by maximizing predicted information gain over the belief space. We estimate information gain using a lightweight LLM-based surrogate and assess world alignment through a novel reward that quantifies the consistency between posterior belief and ground-truth environment configuration. Experiments show that our method outperforms inference-time scaling baselines such as prompt-augmented or retrieval-enhanced LLMs, in aligning with latent world states with significantly lower integration overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
相关 Paper
- Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy OptimizationXingyuan Hua, Sheng Yue, Ju RenICML 2026 · 被引用 1 次
- AdaMEM: Test-Time Adaptive Memory for Language AgentsYunxiang Zhang, Yiheng Li, Ali Payani, Lu WangICML 2026
- Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information ForagingHongjin Qian, Zheng LiuNeurIPS 2025 · 被引用 24 次
- Provable and Practical In-Context Policy Optimization for Self-ImprovementTianrun Yu, yuxiao Yang, Zhaoyang Wang, Kaixiang Zhao 等ICLR 2026 · 被引用 1 次
- Test-Time Adaptation for LLM Agents via Environment InteractionArthur Chen, Zuxin Liu, Jianguo Zhang, Akshara Prabhakar 等ICLR 2026 · 被引用 18 次
