Pixel-Aligned Recurrent Queries for Multi-View 3D Object Detection
Yiming Xie, Huaizu Jiang, Georgia Gkioxari, Julian Straub
摘要
We present PARQ - a multi-view 3D object detector with transformer and pixel-aligned recurrent queries. Unlike previous works that use learnable features or only encode 3D point positions as queries in the decoder, PARQ leverages appearance-enhanced queries initialized from reference points in 3D space and updates their 3D location with recurrent cross-attention operations. Incorporating pixel-aligned features and cross attention enables the model to encode the necessary 3D-to-2D correspondences and capture global contextual information of the input images. PARQ outperforms prior best methods on the ScanNet and ARKitScenes datasets, learns and detects faster, is more robust to distribution shifts in reference points, can leverage additional input views without retraining, and can adapt inference compute by changing the number of recurrent iterations. Code is available at https://ymingxie.github.io/parq.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Redundant Queries in DETR-Based 3D Detection Methods: Unnecessary and PrunableLizhen Xu, Zehao Wu, Wenzhao Qiu, Shanmin Pang 等AAAI 2026 · 被引用 6 次
- DaReNeRF: Direction-aware Representation for Dynamic ScenesAnge Lou, Benjamin Planche, Zhongpai Gao, Yamin Li 等CVPR 2024 · 被引用 4 次
- Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionRunmin Zhang, Zhu Yu, Si-Yuan Cao, Lingyu Zhu 等ICCV 2025 · 被引用 3 次
- VisDiff: SDF-Guided Polygon Generation for Visibility Reconstruction, Characterization and RecognitionRahul Moorthy Mahesh, Jun-Jee Chao, Volkan IslerNeurIPS 2025 · 被引用 1 次
- GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object DetectorZechuan Li, Hongshan Yu, Yihao Ding, Jinhao Qiao 等CVPR 2025
它引用的顶会 Paper11
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima 等ICCV 2019 · 被引用 1,411 次
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li 等ICCV 2023 · 被引用 513 次
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin 等ICCV 2019 · 被引用 242 次
- NeRF-Det: Learning Geometry-Aware Volumetric Representation for Multi-View 3D Object DetectionChenfeng Xu, Bichen Wu, Ji Hou, Sam S. Tsai 等ICCV 2023 · 被引用 71 次
- PlanarRecon: Realtime 3D Plane Detection and Reconstruction from Posed Monocular VideosYiming Xie, Matheus Gadelha, Fengting Yang, Xiaowei Zhou 等CVPR 2022 · 被引用 32 次
相关 Paper
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich 等AAAI 2025 · 被引用 7 次
- 3D Object Detection With PointformerXuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li 等CVPR 2021
- Uni3DETR: Unified 3D Detection TransformerZhenyu Wang, Ya-Li Li, Xi Chen, Hengshuang Zhao 等NeurIPS 2023 · 被引用 65 次
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object DetectionYang Cao, Feize Wu, Dave Chen, Yingji Zhong 等CVPR 2026 · 被引用 6 次
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 被引用 602 次
