RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
Isaac Robinson, Peter Robicheaux, Matvei Popov, Deva Ramanan, Neehar Peri
摘要
Open-vocabulary detectors achieve impressive performance on COCO, but often fail to generalize to real-world datasets with out-of-distribution classes not typically found in their pre-training. Rather than simply fine-tuning a heavy-weight vision-language model (VLM) for new domains, we introduce RF-DETR, a light-weight specialist detection transformer that discovers accuracy-latency Pareto curves for any target dataset with weight-sharing neural architecture search (NAS). Our approach fine-tunes a pre-trained base network on a target dataset and evaluates thousands of network configurations with different accuracy-latency tradeoffs without re-training. Further, we revisit the"tunable knobs"for NAS to improve the transferability of DETRs to diverse target domains. Notably, RF-DETR significantly improves over prior state-of-the-art real-time methods on COCO and Roboflow100-VL. RF-DETR (nano) achieves 48.0 AP on COCO, beating D-FINE (nano) by 5.3 AP at similar latency, and RF-DETR (2x-large) outperforms GroundingDINO (tiny) by 1.2 AP on Roboflow100-VL while running 20x as fast. To the best of our knowledge, RF-DETR (2x-large) is the first real-time detector to surpass 60 AP on COCO. Our code is available at https://github.com/roboflow/rf-detr
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder HelpsXuanlong Yu, Youyang Sha, Longfei Liu, Xi Shen 等CVPR 2026 · 被引用 3 次
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question AnsweringDan Ben Ami, Gabriele Serussi, Kobi Cohen, Chaim BaskinCVPR 2026 · 被引用 3 次
- EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformerMunish Monga, Vishal Chudasama, Pankaj Wasnik, C.V. JawaharCVPR 2026 · 被引用 1 次
- Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object DetectionJialei Zhan, Li Liu, Jiehua Zhang, Yuhang Xie 等CVPR 2026
- MessToClean: Evidence-Grounded Structure-Preserving Reconstruction for Real-World Degraded Exam Paper ImagesJiayi Tuo, Cheng Tang, Zihan Wang, Chenyue Zhou 等ACL 2026
它引用的顶会 Paper15
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 被引用 2,075 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
相关 Paper
- Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-Time Open-Vocabulary Object DetectionYehao Lu, Minghe Weng, Zekang Xiao, Rui Jiang 等ICCV 2025 · 被引用 2 次
- Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object DetectionChuhan Zhang, Chaoyang Zhu, Pingcheng Dong, Long Chen 等ICLR 2025
- Exploring Region-Word Alignment in Built-in Detector for Open-Vocabulary Object DetectionHeng Zhang, Qiuyu Zhao, Linyu Zheng, Hao Zeng 等CVPR 2024 · 被引用 6 次
- Comprehensive Multi-Modal Prototypes Are Simple and Effective Classifiers for Vast-Vocabulary Object DetectionYitong Chen, Wenhao Yao, Lingchen Meng, Sihong Wu 等AAAI 2025
- D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution RefinementYansong Peng, Hebei Li, Peixi Wu, Yueyi Zhang 等ICLR 2025
