Superpowering Open-Vocabulary Object Detectors for X-ray Vision
Pablo Garcia-Fernandez, Lorenzo Vaquero, Mingxuan Liu, Feng Xue, Daniel Cores, Nicu Sebe, Manuel Mucientes, Elisa Ricci
摘要
Open-vocabulary object detection (OvOD) is set to revolutionize security screening by enabling systems to recognize any item in X-ray scans. However, developing effective OvOD models for -ray imaging presents unique challenges due to data scarcity and the modality gap that prevents direct adoption of RGB-based solutions. To overcome these limitations, we propose , a training-free framework that repurposes off-the-shelf RGB OvOD detectors for robust X-ray detection. RAXO builds high-quality X-ray class descriptors using a dual-source retrieval strategy. It gathers relevant RGB images from the web and enriches them via a novel X-ray material transfer mechanism, eliminating the need for labeled databases. These visual descriptors replace text-based classification in OvOD, leveraging intra-modal feature distances for robust detection. Extensive experiments demonstrate that RAXO consistently improves OvOD performance, providing an average increase of up to 17.0 points over base detectors. To further support research in this emerging field, we also introduce DET-COMPASS, a new benchmark featuring bounding box annotations for over 300 object categories, enabling large-scale evaluation of OvOD in X-ray. Code and dataset available at: https://pagf188.github.io/RAXO/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world DetectionLewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang 等NeurIPS 2022 · 被引用 285 次
- Occluded Prohibited Items Detection: An X-ray Security Inspection Benchmark and De-occlusion Attention ModuleYanlu Wei, Renshuai Tao, Zhangjie Wu, Yuqing Ma 等ACM MM 2020 · 被引用 264 次
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICLR 2024 · 被引用 129 次
相关 Paper
- Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item DetectionChuang Peng, Renshuai Tao, Zhongwei Ren, Xianglong Liu 等CVPR 2026 · 被引用 1 次
- Training an Open-Vocabulary Monocular 3D Detection Model without 3D DataRui Huang, Henry Zheng, Yan Wang, Zhuofan Xia 等NeurIPS 2024 · 被引用 26 次
- Zoo3D: Zero-Shot 3D Object Detection at Scene LevelAndrey Lemeshko, Bulat Gabdullin, Nikita Drozdov, Anton Konushin 等CVPR 2026 · 被引用 5 次
- ImOV3D: Learning Open Vocabulary Point Clouds 3D Object Detection from Only 2D ImagesTiming Yang, Yuanliang Ju, Li YiNeurIPS 2024 · 被引用 22 次
- LOVD: Large-and-Open Vocabulary Object DetectionShiyu Tang, Zhaofan Luo, Yifan Wang, Lijun Wang 等ACM MM 2024 · 被引用 1 次
