Fine-Grained Perception in Panoramic Scenes: A Novel Task, Dataset, and Method for Object Importance Ranking
Jia Song, Chenglizhao Chen, Xu Yu, Shanchen Pang
Abstract
Existing Salient Object Ranking (SOR) aims to infer ranking of salient objects based on their saliency degree. However, it tends to only focus on salient objects while neglecting non-salient ones. This coarse-grained ranking limits the performance of downstream tasks. For instance, in image retrieval tasks, focusing solely on the relationship between salient objects is insufficient for achieving fine-grained scene analysis, which may result in retrieved results that do not satisfy user requirements. High-quality retrieval requires fine-grained analysis, making it essential to rank non-salient objects. Based on this need, we propose a new task: Fine-grained Object Importance Ranking in 360 Scenes (FOIR-360), which focus on predicting the relative importance of "ALL objects'' at the instance-level. Our task takes into account all objects, allowing us to refine the original "coarse-grained'' to a "fine-grained'' level. Currently, the main challenge for this new task is the lack of supervised data for model training or even for model testing. Therefore, we propose a novel weakly supervised method to address the shortage of datasets. Furthermore, to the best of our knowledge, there is no existing suitable annotation protocol for this new task. The main reason is that annotating fine-grained rankings is extremely difficult, especially in panoramic scenes that contain numerous instances where even humans are unable to determine which one is more important than others. As the first attempt, we introduce a new annotation protocol designed to highlight the ranking of objects that are non-salient yet still important. Based on this protocol, we construct the first fine-grained 360Rank dataset. In summary, all these new task, weakly supervised method, annotation protocol, and dataset have the potential to drive advancements in the field.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi et al.CVPR 2022 · 382 citations
- Bi-directional Object-Context Prioritization Learning for Saliency RankingXin Tian, Ke Xu, Xin Yang, Lin Du et al.CVPR 2022 · 33 citations
- SeqRank: Sequential Ranking of Salient ObjectsHuankang Guan, Rynson W. H. LauAAAI 2024 · 7 citations
- Partitioned Saliency Ranking with Dense Pyramid TransformersChengxiao Sun, Yan Xu, Jialun Pei, Haopeng Fang et al.ACM MM 2023 · 4 citations
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
Related papers
- Salient Object Ranking with Position-Preserved AttentionHao Fang, Daoxin Zhang, Yi Zhang, Minghao Chen et al.ICCV 2021 · 26 citations
- Instance-Level Panoramic Audio-Visual Saliency Detection and RankingRuohao Guo, Dantong Niu, Liao Qu, Yanyu Qi et al.ACM MM 2024
- Probabilistic Salient Object RankingRongjin Guo, Guan Huankang, Rynson W LauICML 2026
- Salient Object Ranking via Cyclical Perception-Viewing Interaction ModelingRongjin Guo, Ke Xu, Rynson W. H. LauICLR 2026
- Advancing Saliency Ranking with Human Fixations: Dataset, Models and BenchmarksBowen Deng, Siyang Song, Andrew P. French, Denis Schluppeck et al.CVPR 2024
