STRNet: Visual Navigation with Spatio-Temporal Representation through Dynamic Graph Aggregation
Hao Ren, Zetong Bi, Yiming Zeng, Zhaoliang Wan, Lu Qi, Hui Cheng
Abstract
Visual navigation requires the robot to reach a specified goal such as an image, based on a sequence of first-person visual observations. While recent learning-based approaches have made significant progress, they often focus on improving policy heads or decision strategies while relying on simplistic feature encoders and temporal pooling to represent visual input. This leads to the loss of fine-grained spatial and temporal structure, ultimately limiting accurate action prediction and progress estimation. In this paper, we propose a unified spatio-temporal representation framework that enhances visual encoding for robotic navigation. Our approach extracts features from both image sequences and goal observations, and fuses them using the designed spatio-temporal fusion module. This module performs spatial graph reasoning within each frame and models temporal dynamics using a hybrid temporal shift module combined with multi-resolution difference-aware convolution. Experimental results demonstrate that our approach consistently improves navigation performance and offers a generalizable visual backbone for goal-conditioned control. Code is available at https://github.com/hren20/STRNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e7d265d-a112-4868-a089-b1145700e974Cited by top-tier papers1
Ask how each one uses itBuilds on12
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman et al.NeurIPS 2022 · 344 citations
- Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual NavigationZiad Al-Halah, Santhosh K. Ramakrishnan, Kristen GraumanCVPR 2022 · 52 citations
Related papers
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt et al.ICCV 2023 · 56 citations
- VTNet: Visual Transformer Network for Object Goal NavigationHeming Du, Xin Yu, Liang ZhengICLR 2021 · 36 citations
- FGPrompt: Fine-grained Goal Prompting for Image-goal NavigationXinyu Sun, Peihao Chen, Jugang Fan, Jian Chen et al.NeurIPS 2023 · 41 citations
- NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object NavigationWei Xie, Haobo Jiang, Yun Zhu, Jianjun Qian et al.AAAI 2025 · 7 citations
- End-to-End (Instance)-Image Goal Navigation through Correspondence as an Emergent PhenomenonGuillaume Bono, Leonid Antsfeld, Boris Chidlovskii, Philippe Weinzaepfel et al.ICLR 2024 · 19 citations
