Causal Abstraction Learning for Multi-Modal Grounded Planning
Xinshu Li, Shiyi Yang, Ziqi Xu, Feng Xia, Quan Z. Sheng, Lina Yao
摘要
Recent advances in multimodal embodied agents have enabled long-horizon planning in visually rich environments via natural language. Yet, their generalization remains brittle when task instructions deviate from familiar examples, exposing a reliance on surface imitation rather than structural understanding. We propose Causal Abstraction Learning for Multi-Modal Grounded Planning (CALM), a framework that enhances planning agents with the ability to discover and exploit causal regularities across tasks. CALM incrementally develops a causal library by abstracting precondition–effect structure from successful executions, yielding compact representations that emphasize stable dependencies beyond incidental context. When execution diverges from expectation, these abstractions are refined through contrastive causal reasoning, enabling targeted adjustments that resolve underlying mechanism mismatch. The resulting structure serves as a transferable prior for planning in novel settings, integrating perceptual cues with mechanism-informed knowledge. Without retraining or task-specific heuristics, CALM generalizes robustly and efficiently to linguistic and perceptual variation. Experiments on ALFRED and VirtualHome demonstrate consistent gains, highlighting causal abstraction as a scalable inductive bias for grounded planning.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot PlanningYichao Liang, Dat Nguyen, Cambridge Yang, Tianyang Li 等ICLR 2026 · 被引用 11 次
- Learning Grounded Action Abstractions from LanguageLionel Wong, Jiayuan Mao, Pratyusha Sharma, Zachary S. Siegel 等ICLR 2024 · 被引用 7 次
- Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction FollowingMinjong Yoo, Jinwoo Jang, Wei-Jin Park, Honguk WooNeurIPS 2024 · 被引用 15 次
- Language Agents Meet Causality - Bridging LLMs and Causal World ModelsJohn Gkountouras, Matthias Lindemann, Phillip Lippe, Efstratios Gavves 等ICLR 2025
- Egocentric Planning for Scalable Embodied Task AchievementXiaotian Liu, Héctor Palacios, Christian MuiseNeurIPS 2023 · 被引用 9 次
