Fine-grained Zero-Shot Object Detection
Hongxu Ma, Chenbo Zhang, Lu Zhang, Jiaogen Zhou, Jihong Guan, Shuigeng Zhou
摘要
Zero-shot object detection (ZSD) aims to leverage semantic descriptions to localize and recognize objects of both seen and unseen classes. Existing ZSD works are mainly coarse-grained object detection, where the classes are visually quite different, thus are relatively easy to distinguish. However, in real life we often have to face fine-grained object detection scenarios, where the classes are too similar to be easily distinguished. For example, detecting different kinds of birds, fishes, and flowers. In this paper, we propose and solve a new problem called Fine-Grained Zero-Shot Object Detection (FG-ZSD for short), which aims to detect objects of different classes with minute differences in details under the ZSD paradigm. We develop an effective method called MSHC for the FG-ZSD task, which is based on an improved two-stage detector and employs a multi-level semantics-aware embedding alignment loss, ensuring tight coupling between the visual and semantic spaces. Considering that existing ZSD datasets are not suitable for the new FG-ZSD task, we build the first FG-ZSD benchmark dataset FGZSD-Birds, which contains 148,820 images falling into 36 orders, 140 families, 579 genera and 1432 species. Extensive experiments on FGZSD-Birds show that our method outperforms existing ZSD models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- RAM: Recover Any 3D Human Motion in-the-WildSen Jia, Ning Zhu, Jinqin Zhong, Jiale Zhou 等CVPR 2026 · 被引用 12 次
- MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language ModelsYang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu 等ACL 2026 · 被引用 9 次
- Multimodal LLMs Can Reason about Aesthetics in Zero-ShotRuixiang Jiang, Chang Wen ChenACM MM 2025 · 被引用 6 次
- Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video SynthesisYanzuo Lu, Yuxi Ren, Xin Xia, Shanchuan Lin 等ICCV 2025 · 被引用 4 次
- RL-ScanIQA: Reinforcement-Learned Scanpaths for Blind 360deg Image Quality AssessmentYujia Wang, Yuyan Li, Jiuming Liu, Fang-Lue Zhang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper22
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing 等CVPR 2022 · 被引用 296 次
- HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot LearningShiming Chen, Guo-Sen Xie, Yang Liu, Qinmu Peng 等NeurIPS 2021 · 被引用 190 次
- Few-Shot Object Detection with Fully Cross-TransformerGuangxing Han, Jiawei Ma, Shiyuan Huang, Long Chen 等CVPR 2022 · 被引用 183 次
相关 Paper
- Zero-Shot Object Detection by Semantics-Aware DETR with Adaptive Contrastive LossHuan Liu, Lu Zhang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 被引用 6 次
- Transductive Learning for Zero-Shot Object DetectionShafin Rahman, Salman H. Khan, Nick BarnesICCV 2019 · 被引用 82 次
- Meta-ZSDETR: Zero-shot DETR with Meta-learningLu Zhang, Chenbo Zhang, Jiajia Zhao, Jihong Guan 等ICCV 2023 · 被引用 10 次
- Hierarchical Few-Shot Object Detection: Problem, Benchmark and MethodLu Zhang, Yang Wang, Jiaogen Zhou, Chenbo Zhang 等ACM MM 2022 · 被引用 13 次
- Learning the Redundancy-Free Features for Generalized Zero-Shot Object RecognitionZongyan Han, Zhenyong Fu, Jian YangCVPR 2020
