Towards Content-based Pixel Retrieval in Revisited Oxford and Paris
Guoyuan An, Woo Jae Kim, Saelyne Yang, Rong Li, Yuchi Huo, Sung-Eui Yoon
Abstract
This paper introduces the first two landmark pixel retrieval benchmarks. Pixel retrieval is segmented instance retrieval. Like semantic segmentation extends classification to the pixel level, pixel retrieval is an extension of image retrieval and offers information about which pixels are related to the query object. In addition to retrieving images for the given query, it helps users quickly identify the query object in true positive images and exclude false positive images by denoting the correlated pixels. Our user study results show pixel-level annotation can significantly improve the user experience. Compared with semantic and instance segmentation, pixel retrieval requires a fine-grained recognition capability for variable-granularity targets. To this end, we propose pixel retrieval benchmarks named PROxford and PRParis, which are based on the widely used image retrieval datasets, ROxford and RParis. Three professional annotators label 5,942 images with two rounds of double-checking and refinement. Furthermore, we conduct extensive experiments and analysis on the SOTA methods in image search, image matching, detection, segmentation, and dense matching using our pixel retrieval benchmarks. Results show that the pixel retrieval task is challenging to these approaches and distinctive from existing problems, suggesting that further research can advance the content-based pixel-retrieval and thus user search experience. The datasets can be downloaded from this link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on12
- YOLACT: Real-Time Instance SegmentationDaniel Bolya, Chong Zhou, Fanyi Xiao, Yong Jae LeeICCV 2019 · 2,075 citations
- LiT: Zero-Shot Transfer with Locked-image text TuningXiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner et al.CVPR 2022 · 349 citations
- Mining Latent Classes for Few-shot SegmentationLihe Yang, Wei Zhuo, Lei Qi, Yinghuan Shi et al.ICCV 2021 · 152 citations
- GOCor: Bringing Globally Optimized Correspondence Volumes into Your Neural NetworkPrune Truong, Martin Danelljan, Luc Van Gool, Radu TimofteNeurIPS 2020 · 89 citations
- Beyond Cross-view Image Retrieval: Highly Accurate Vehicle Localization Using Satellite ImageYujiao Shi, Hongdong LiCVPR 2022 · 81 citations
Related papers
- Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and RetrievalTobias Weyand, André Araújo, Bingyi Cao, Jack SimCVPR 2020
- Rethinking Benchmarks for Cross-modal Image-text RetrievalWeijing Chen, Linli Yao, Qin JinSIGIR 2023 · 25 citations
- UFineBench: Towards Text-based Person Retrieval with Ultra-fine GranularityJialong Zuo, Hanyu Zhou, Ying Nie, Feng Zhang et al.CVPR 2024 · 45 citations
- Instance-level Image Retrieval using Reranking TransformersFuwen Tan, Jiangbo Yuan, Vicente OrdonezICCV 2021 · 116 citations
- OmniCity: Omnipotent City Understanding with Multi-Level and Multi-View ImagesWeijia Li, Yawen Lai, Linning Xu, Yuanbo Xiangli et al.CVPR 2023
