Continuous Spatiotemporal Transformer
Antonio Henrique de Oliveira Fonseca, Emanuele Zappala, Josue Ortega Caro, David van Dijk
摘要
To enhance the obstacle avoidance performance and autonomous decision-making capabilities of robots in complex dynamic environments, this paper proposes an end-to-end intelligent obstacle avoidance method that integrates deep reinforcement learning, spatiotemporal attention mechanisms, and a Transformer-based architecture. Current mainstream robot obstacle avoidance methods often rely on system architectures with separated perception and decision-making modules, which suffer from issues such as fragmented feature transmission, insufficient environmental modeling, and weak policy generalization. To address these problems, this paper adopts Deep Q-Network (DQN) as the core of reinforcement learning, guiding the robot to autonomously learn optimal obstacle avoidance strategies through interaction with the environment, effectively handling continuous decision-making problems in dynamic and uncertain scenarios. To overcome the limitations of traditional perception mechanisms in modeling the temporal evolution of obstacles, a spatiotemporal attention mechanism is introduced, jointly modeling spatial positional relationships and historical motion trajectories to enhance the model's perception of critical obstacle areas and potential collision risks. Furthermore, an end-to-end Transformer-based perception-decision architecture is designed, utilizing multi-head self-attention to perform high-dimensional feature modeling on multi-modal input information (such as LiDAR and depth images), and generating action policies through a decoding module. This completely eliminates the need for manual feature engineering and intermediate state modeling, constructing an integrated learning process of perception and decision-making. Experiments conducted in several typical obstacle avoidance simulation environments demonstrate that the proposed method outperforms existing mainstream deep reinforcement learning approaches in terms of obstacle avoidance success rate, path optimization, and policy convergence speed. It exhibits good stability and generalization capabilities, showing broad application prospects for deployment in real-world complex environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Theoretical Understandings of Self-Consuming Generative ModelsShi Fu, Sen Zhang, Yingjie Wang, Xinmei Tian 等ICML 2024 · 被引用 26 次
- Spectral-Refiner: Accurate Fine-Tuning of Spatiotemporal Fourier Neural Operator for Turbulent FlowsShuhao Cao, Francesco Brarda, Ruipeng Li, Yuanzhe XiICLR 2025 · 被引用 1 次
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal TransformersRuihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu 等ICLR 2022 · 被引用 146 次
- DiffE2E: Rethinking End-to-End Driving with a Hybrid Diffusion-Regression-Classification PolicyRui Zhao, Yuze Fan, Ziguo Chen, Fei Gao 等NeurIPS 2025 · 被引用 7 次
- Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery ProblemsZengyu Zou, Jingyuan Wang, Yixuan Huang, Junjie WuAAAI 2026
- A Consciousness-Inspired Planning Agent for Model-Based Reinforcement LearningMingde Zhao, Zhen Liu, Sitao Luan, Shuyuan Zhang 等NeurIPS 2021 · 被引用 41 次
- Parameterized Decision-Making with Multi-Modality Perception for Autonomous DrivingYuyang Xia, Shuncheng Liu, Quanlin Yu, Liwei Deng 等ICDE 2024 · 被引用 26 次
