Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
Nhat Chung, Taisei Hanyu, Toan Nguyen, Huy Le, Frederick Bumgarner, Duy Minh Ho Nguyen, Khoa Vo, Kashu Yamazaki, Chase Rainwater, Tung Kieu, Anh Nguyen, Ngan Le
摘要
As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In these non-Markovian settings, key decision cues are often hidden in object-specific histories rather than the current scene. Without persistent memory of prior interactions (what has been interacted with, where it has been, or how it has changed) visuomotor policies may fail, repeat past actions, or overlook completed ones. To surface this challenge, we introduce LIBERO-Mem, a non-Markovian task suite for stress-testing robotic manipulation under object-level partial observability. It combines short-and long-horizon object tracking with temporally sequenced subgoals, requiring reasoning beyond the current frame. However, vision-language-action (VLA) models often struggle in such settings, with token scaling quickly becoming intractable even for tasks spanning just a few hundred frames. We propose Embodied-SlotSSM, a slotcentric VLA framework built for temporal scalability. It maintains spatio-temporally consistent slot identities and leverages them through two mechanisms: (1) slot-state-space modeling for reconstructing short-term history, and (2) a relational encoder to align the input tokens with action decoding. Together, these components enable temporally grounded, context-aware action prediction. Experiments show Embodied-SlotSSM's baseline performance on LIBERO-Mem and general tasks, offering a scalable solution for non-Markovian reasoning in object-centric robotic policies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action AdaptationDuc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho 等ICML 2026 · 被引用 3 次
- AffordMatcher: Affordance Learning in 3D Scenes from Visual SignifiersNghia Vu, Tuong Do, Khang Nguyen, Baoru Huang 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper16
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon 等CVPR 2022 · 被引用 398 次
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani 等ICLR 2021 · 被引用 357 次
相关 Paper
- HAMLET: Switch Your Vision-Language-Action Model into a History-Aware PolicyMyungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee 等ICLR 2026 · 被引用 52 次
- AVA-VLA: Improving Vision-Language-Action models with Active Visual AttentionLei Xiao, Jifeng Li, Juntao Gao, Feiyang Ye 等CVPR 2026 · 被引用 26 次
- Spatial Memory for Out-of-Vision Manipulation in Vision-Language-ActionPengteng Li, Weiyu Guo, He ZHANG, Tiefu Cai 等ICML 2026 · 被引用 3 次
- Reusable Slotwise MechanismsBailey Trang Nguyen, Amin Mansouri, Kanika Madan, Khuong Nguyen 等NeurIPS 2023 · 被引用 6 次
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic ManipulationHao Shi, Bin Xie, Yingfei Liu, Lin Sun 等ICLR 2026 · 被引用 227 次
