Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement
Xiuquan Hou, Meiqin Liu, Senlin Zhang, Ping Wei, Badong Chen
Abstract
DETR-like methods have significantly increased detection performance in an end-to-end manner. The mainstream two-stage frameworks of them perform dense selfattention and select a fraction of queries for sparse crossattention, which is proven effective for improving performance but also introduces a heavy computational burden and high dependence on stable query selection. This paper demonstrates that suboptimal two-stage selection strategies result in scale bias and redundancy due to the mismatch between selected queries and objects in two-stage initialization. To address these issues, we propose hierarchical salience filtering refinement, which performs transformer encoding only on filtered discriminative queries, for a better trade-off between computational efficiency and precision. The filtering process overcomes scale bias through a novel scale-independent salience supervision. To compensate for the semantic misalignment among queries, we introduce elaborate query refinement modules for stable two-stage initialization. Based on above improvements, the proposed Salience DETR achieves significant improvements of +4.0% AP, +0.2% AP, +4.4% AP on three challenging task-specific detection datasets, as well as 49.2% AP on
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image GenerationYunhong Min, Daehyeon Choi, Kyeongmin Yeo, Jihyun Lee et al.NeurIPS 2025 · 10 citations
- PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object DetectionZhengjian Kang, Jun Zhuang, Kangtong Mo, Qi Chen et al.CVPR 2026 · 6 citations
- LMM-Det: Make Large Multimodal Models Excel in Object DetectionJincheng Li, Chunyu Xie, Ji Ao, Dawei Leng et al.ICCV 2025 · 2 citations
- Fractional Correspondence Framework in Detection TransformerMasoumeh Zareapoor, Pourya Shamsolmoali, Huiyu Zhou, Yue Lu et al.ACM MM 2024 · 1 citation
- OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware MiningBingzhi Chen, Sisi Fu, Xiaocheng Fang, Jieyi Cai et al.CVPR 2025
Builds on16
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- TOOD: Task-aligned One-stage Object DetectionChengjian Feng, Yujie Zhong, Yu Gao, Matthew R. Scott et al.ICCV 2021 · 1,191 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
Related papers
- Sparse Semi-DETR: Sparse Learnable Queries for Semi-Supervised Object DetectionTahira Shehzadi, Khurram Azeem Hashmi, Didier Stricker, Muhammad Zeshan AfzalCVPR 2024 · 36 citations
- Decoupled DETR: Spatially Disentangling Localization and Classification for Improved End-to-End Object DetectionManyuan Zhang, Guanglu Song, Yu Liu, Hongsheng LiICCV 2023 · 31 citations
- Sparse DETR: Efficient End-to-End Object Detection with Learnable SparsityByungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon KimICLR 2022 · 256 citations
- Less is More: Focus Attention for Efficient DETRDehua Zheng, Wenhui Dong, Hailin Hu, Xinghao Chen et al.ICCV 2023 · 128 citations
- Semi-DETR: Semi-Supervised Object Detection with Detection TransformersJiacheng Zhang, Xiangru Lin, Wei Zhang, Kuo Wang et al.CVPR 2023
