PRISM: Perception Reasoning Interleaved for Sequential Decision Making.
Mohamed Salim AISSI, Salim Aissi, Clément Romac, Laure Soulier, Mohamed Chetouani, Olivier Sigaud, Nicolas THOME
摘要
Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception–reasoning–decision gap in standalone Vision–Language Models (VLMs), which often overlook task-critical information. In this paper, we introduce PRISM, a framework that tightly couples perception (VLM) and decision (LLM) through a dynamic question–answer (DQA) pipeline. Instead of passively accepting the VLM’s description, the LLM critiques it, probes the VLM with goal-oriented questions, and synthesizes a compact image description. This closed-loop interaction yields a sharp, task-driven understanding of the scene. We evaluate PRISM on the ALFWorld and Room-to-Room (R2R) benchmarks. We show that: (1) PRISM significantly outperforms state-of-the-art image-based models, (2) our Interactive goal-oriented perception pipeline yields systematic and substantial gains, and (3) PRISM is fully automatic, eliminating the need for handcrafted questions or answers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk 等ICLR 2021 · 被引用 819 次
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
相关 Paper
- Prism: A Framework for Decoupling and Assessing the Capabilities of VLMsYuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang 等NeurIPS 2024 · 被引用 49 次
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical EnvironmentsWeijie Zhou, Xuantang Xiong, Yi Peng, Manli Tao 等NeurIPS 2025 · 被引用 4 次
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao 等CVPR 2024 · 被引用 23 次
- Grounded Semantic Role Labelling from Synthetic Multimodal Data for Situated Robot CommandsClaudiu Daniel Hromei, Antonio Scaiella, Danilo Croce, Roberto BasiliEMNLP 2025 · 被引用 1 次
- ActiView: Evaluating Active Perception Ability for Multimodal Large Language ModelsZiyue Wang, Chi Chen, Fuwen Luo, Yurui Dong 等ACL 2025 · 被引用 9 次
