Kinaema: a recurrent sequence model for memory and pose in motion
Mert Bülent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci, Christian Wolf
Abstract
One key aspect of spatially aware robots is the ability to "find their bearings", i.e. to correctly situate themselves in previously seen spaces. In this work, we focus on this particular scenario of continuous robotics operations, where information observed before an actual episode start is exploited to optimize efficiency. We introduce a new model, Kinaema, and agent, capable of integrating a stream of visual observations while moving in a potentially large scene, and upon request, processing a query image and predicting the relative position of the shown space with respect to its current position. Our model does not explicitly store an observation history, therefore does not have hard constraints on context length. It maintains an implicit latent memory, which is updated by a transformer in a recurrent way, compressing the history of sensor readings into a compact representation. We evaluate the impact of this model in a new downstream task we call "Mem-Nav". We show that our large-capacity recurrent model maintains a useful representation of the scene, navigates to goals observed before the actual episode start, and is computationally efficient, in particular compared to classical transformers with attention over an observation history.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0abdd22-5407-40c0-9d5d-0b3930e70c53Cited by top-tier papers1
Ask how each one uses itBuilds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
Related papers
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal et al.NeurIPS 2025 · 9 citations
- Token Turing MachinesMichael S. Ryoo, Keerthana Gopalakrishnan, Kumara Kahatapitiya, Ted Xiao et al.CVPR 2023
- History Aware Multimodal Transformer for Vision-and-Language NavigationShizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, Ivan LaptevNeurIPS 2021 · 427 citations
- MemoNav: Working Memory Model for Visual NavigationHongxin Li, Zeyu Wang, Xu Yang, Yuran Yang et al.CVPR 2024
- VLN BERT: A Recurrent Vision-and-Language BERT for NavigationYicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo et al.CVPR 2021
