SIRA: Scalable Inter-Frame Relation and Association for Radar Perception
Ryoma Yataka, Pu Wang, Petros Boufounos, Ryuhei Takahashi
Abstract
Conventional radar feature extraction faces limitations due to low spatial resolution, noise, multipath reflection, the presence of ghost targets, and motion blur. Such limitations can be exacerbated by nonlinear object motion, particularly from an ego-centric viewpoint. It becomes evident that to address these challenges, the key lies in exploiting temporal feature relation over an extended horizon and enforcing spatial motion consistency for effective association. To this end, this paper proposes SIRA (Scalable Inter-frame Relation and Association) with two designs. First, inspired by Swin Transformer, we introduce extended temporal relation, generalizing the existing temporal relation layer from two consecutive frames to multiple inter-frames with temporally regrouped window attention for scalability. Second, we propose motion consistency track with the concept of a pseudo-tracklet generated from observational data for better trajectory prediction and subsequent object association. Our approach achieves 58.11 mAP@0.5 for oriented object detection and 47.79 MOTA for multiple object tracking on the Radiate dataset, surpassing previous state-of-the-art by a margin of +4.11 mAP@0.5 and +9.94 MOTA, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9c3a197-cd44-4fe0-91a8-a63637822e15Cited by top-tier papers2
- RETR: Multi-View Radar Detection Transformer for Indoor PerceptionRyoma Yataka, Adriano Cardace, Perry Wang, Petros Boufounos et al.NeurIPS 2024 · 21 citations
- Towards Accurate 3D Object Detection in Adverse Weather by Leveraging 4D Radar for LiDAR Geometry EnhancementTianxu Tong, Xinrun Liu, Hongmin Liu, Bin FanAAAI 2026
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- Raw High-Definition Radar for Multi-Task LearningJulien Rebut, Arthur Ouaknine, Waqas Malik, Patrick PérezCVPR 2022 · 102 citations
- Multi-View Radar Semantic SegmentationArthur Ouaknine, Alasdair Newson, Patrick Pérez, Florence Tupin et al.ICCV 2021 · 98 citations
Related papers
- Exploiting Temporal Relations on Radar Perception for Autonomous DrivingPeizhao Li, Pu Wang, Karl Berntorp, Hongfu LiuCVPR 2022 · 50 citations
- Delving into Motion-Aware Matching for Monocular 3D Object TrackingKuan-Chih Huang, Ming-Hsuan Yang, Yi-Hsuan TsaiICCV 2023 · 20 citations
- MARSS: Radar Semantic Segmentation via Modular Attention and State Space ModelsFengyu Chen, Tiao Tan, Teng Li, Yuantian Quan et al.CVPR 2026
- TARS: Traffic-Aware Radar Scene Flow EstimationJialong Wu, Marco Braun, Dominic Spata, Matthias RottmannICCV 2025
- CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object DetectionJisong Kim, Minjae Seong, Jun Won ChoiNeurIPS 2024 · 27 citations
