Learning Dynamic Query Combinations for Transformer-based Object Detection and Segmentation
Yiming Cui, Linjie Yang, Haichao Yu
Abstract
Transformer-based detection and segmentation methods use a list of learned detection queries to retrieve information from the transformer network and learn to predict the location and category of one specific object from each query. We empirically find that random convex combinations of the learned queries are still good for the corresponding models. We then propose to learn a convex combination with dynamic coefficients based on the high-level semantics of the image. The generated dynamic queries, named modulated queries, better capture the prior of object locations and categories in the different images. Equipped with our modulated queries, a wide range of DETR-based models achieve consistent and superior performance across multiple tasks including object detection, instance segmentation, panoptic segmentation, and video instance segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a35aba09-cf68-45f1-a298-63df1fbc7cb2Cited by top-tier papers4
- ClusterFomer: Clustering As A Universal Visual LearnerJames Liang, Yiming Cui, Qifan Wang, Tong Geng et al.NeurIPS 2023 · 63 citations
- Dynamic Dictionary Learning for Remote Sensing Image SegmentationXuechao Zou, Yue Li, Shun Zhang, Kai Li et al.ICCV 2025 · 15 citations
- PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object DetectionZhengjian Kang, Jun Zhuang, Kangtong Mo, Qi Chen et al.CVPR 2026 · 6 citations
- Mr. DETR: Instructive Multi-Route Training for Detection TransformersChang-Bin Zhang, Yujie Zhong, Kai HanCVPR 2025
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
Related papers
- UP-DETR: Unsupervised Pre-Training for Object Detection With TransformersZhigang Dai, Bolun Cai, Yugeng Lin, Junying ChenCVPR 2021
- DECO: Unleashing the Potential of ConvNets for Query-based Detection and SegmentationXinghao Chen, Siwei Li, Yijing Yang, Yunhe WangICLR 2025
- SOLQ: Segmenting Objects by Learning QueriesBin Dong, Fangao Zeng, Tiancai Wang, Xiangyu Zhang et al.NeurIPS 2021 · 143 citations
- Feature Aggregated Queries for Transformer-Based Video Object DetectorsYiming CuiCVPR 2023
- Decoupling Dense Video Captioning via Task-specific PromptsWei Chen, Jianwei Niu, Xuefeng Liu, Xinghao WuACM MM 2025 · 1 citation
