RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric Strategies
Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Yao Li, Yanyong Zhang
Abstract
The recent advances in query-based multi-camera 3D object detection are featured by initializing object queries in the 3D space, and then sampling features from perspective-view images to perform multi-round query refinement. In such a framework, query points near the same camera ray are likely to sample similar features from very close pixels, resulting in ambiguous query features and degraded detection accuracy. To this end, we introduce RayFormer, a camera-ray-inspired query-based 3D object detector that aligns the initialization and feature extraction of object queries with the optical characteristics of cameras. Specifically, RayFormer transforms perspective-view image features into bird's eye view (BEV) via the lift-splat-shoot method and segments the BEV map to sectors based on the camera rays. Object queries are uniformly and sparsely initialized along each camera ray, facilitating the projection of different queries onto different areas in the image to extract distinct features. Besides, we leverage the instance information of images to supplement the uniformly initialized object queries by further involving additional queries along the ray from 2D object detection boxes. To extract unique object-level features that cater to distinct queries, we design a ray sampling method that suitably organizes the distribution of feature sampling points on both images and bird's eye view. Extensive experiments are conducted on the nuScenes dataset to validate our proposed ray-inspired model design. The proposed RayFormer achieves superior performance of 55.5% mAP and 63.3% NDS, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object DetectionJiangyong Yu, Changyong Shu, Sifan Zhou, Zichen Yu et al.AAAI 2026
- Seeing in Double: Dual-Granularity BEV Segmentation via Mamba-Driven Alignment and Polar-Decoupled ExpertsJiaxin Cai, Rui Lin, Jingze Su, Qi Li et al.AAAI 2026
- RaCFormer: Towards High-Quality 3D Object Detection via Query-based Radar-Camera FusionXiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan et al.CVPR 2025
- RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object DetectionRui Ding, Zhaonian Kuang, Zongwei Zhou, Meng Yang et al.AAAI 2026
Builds on20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li et al.ICCV 2023 · 513 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
Related papers
- FrustumFormer: Adaptive Instance-aware Resampling for Multi-view 3D DetectionYuqi Wang, Yuntao Chen, Zhaoxiang ZhangCVPR 2023
- Enhancing 3D Object Detection with 2D Detection-Guided Query AnchorsHaoxuanye Ji, Pengpeng Liang, Erkang ChengCVPR 2024
- PolarFormer: Multi-Camera 3D Object Detection with Polar TransformerYanqin Jiang, Li Zhang, Zhenwei Miao, Xiatian Zhu et al.AAAI 2023 · 240 citations
- SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosHaisong Liu, Yao Teng, Tao Lu, Haiguang Wang et al.ICCV 2023 · 204 citations
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang et al.ICLR 2023 · 28 citations
