RETR: Multi-View Radar Detection Transformer for Indoor Perception
Ryoma Yataka, Adriano Cardace, Perry Wang, Petros Boufounos, Ryuhei Takahashi
Abstract
Indoor radar perception has seen rising interest due to affordable costs driven by emerging automotive imaging radar developments and the benefits of reduced privacy concerns and reliability under hazardous conditions (e.g., fire and smoke). However, existing radar perception pipelines fail to account for distinctive characteristics of the multi-view radar setting. In this paper, we propose Radar dEtection TRansformer (RETR), an extension of the popular DETR architecture, tailored for multi-view radar perception. RETR inherits the advantages of DETR, eliminating the need for hand-crafted components for object detection and segmentation in the image plane. More importantly, RETR incorporates carefully designed modifications such as 1) depth-prioritized feature similarity via a tunable positional encoding (TPE); 2) a tri-plane loss from both radar and camera coordinates; and 3) a learnable radar-to-camera transformation via reparameterization, to account for the unique multi-view radar setting. Evaluated on two indoor radar perception datasets, our approach outperforms existing state-of-the-art methods by a margin of 15.38+ AP for object detection and 11.91+ IoU for instance segmentation, respectively. Our implementation is available at https://github.com/merlresearch/radar-detection-transformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh ReconstructionJunqiao Fan, Yunjiao Zhou, Yizhuo Yang, Xinyuan Cui et al.CVPR 2026 · 11 citations
- RISE: Single Static Radar-based Indoor Scene UnderstandingKaichen Zhou, Laura Dodds, Sayed Saad Afzal, Fadel AdibCVPR 2026 · 3 citations
- Person Parametric Physics-informed Representation for mmWave-based Human Pose EstimationShuntian Zheng, Jiaqi Li, Guangming Wang, Minzhe Ni et al.UbiComp 2026 · 1 citation
- Indoor Multi-View Radar Object Detection via 3D Bounding Box DiffusionRyoma Yataka, Pu Perry Wang, Petros Boufounos, Ryuhei TakahashiAAAI 2026 · 1 citation
Builds on13
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- Anchor DETR: Query Design for Transformer-Based DetectorYingming Wang, Xiangyu Zhang, Tong Yang, Jian SunAAAI 2022 · 567 citations
Related papers
- RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object DetectionYiheng Li, Yang Yang, Zhen LeiAAAI 2025 · 4 citations
- RAPTR: Radar-based 3D Pose Estimation using TransformerSorachi Kato, Ryoma Yataka, Pu Perry Wang, Pedro Miraldo et al.NeurIPS 2025 · 5 citations
- Towards Foundational Models for Single-Chip RadarTianshu Huang, Akarsh Prabhakara, Chuhan Chen, Jay Karhade et al.ICCV 2025 · 3 citations
- RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera FusionXiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan et al.CVPR 2025
- RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object DetectionZhiwei Lin, Zhe Liu, Zhongyu Xia, Xinhao Wang et al.CVPR 2024
