TransView: Inside, Outside, and Across the Cropping View Boundaries
Zhiyu Pan, Zhiguo Cao, Kewei Wang, Hao Lu, Weicai Zhong
摘要
We show that relation modeling between visual elements matters in cropping view recommendation. Cropping view recommendation addresses the problem of image recomposition conditioned on the composition quality and the ranking of views (cropped sub-regions). This task is challenging because the visual difference is subtle when a visual element is reserved or removed. Existing methods represent visual elements by extracting region-based convolutional features inside and outside the cropping view boundaries, without probing a fundamental question: why some visual elements are of interest or of discard? In this work, we observe that the relation between different visual elements significantly affects their relative positions to the desired cropping view, and such relation can be characterized by the attraction inside/outside the cropping view boundaries and the repulsion across the boundaries. By instantiating a transformer-based solution that represents visual elements as visual words and that models the dependencies between visual words, we report not only state-of-the-art performance on public benchmarks, but also interesting visualizations that depict the attraction and repulsion between visual elements, which may shed light on what makes for effective cropping view recommendation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Rethinking Image Cropping: Exploring Diverse Compositions from Global ViewsGengyun Jia, Huaibo Huang, Chaoyou Fu, Ran HeCVPR 2022 · 被引用 19 次
- Learning Second-Order Attentive Context for Efficient Correspondence PruningXinyi Ye, Weiyue Zhao, Hao Lu, Zhiguo CaoAAAI 2023 · 被引用 12 次
- Shoot360: Normal View Video Creation from City Panorama FootageAnyi Rao, Linning Xu, Dahua LinSIGGRAPH 2022 · 被引用 5 次
- ProCrop: Learning Aesthetic Image Cropping from Professional CompositionsKe Zhang, Tianyu Ding, Jiachen Jiang, Tianyi Chen 等AAAI 2026 · 被引用 3 次
- Exposure Completing for Temporally Consistent Neural High Dynamic Range Video RenderingJiahao Cui, Wei Jiang, Zhan Peng, Zhiyu Pan 等ACM MM 2024 · 被引用 3 次
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Image Cropping with Composition and Saliency Aware Aesthetic Score MapYi Tu, Li Niu, Weijie Zhao, Dawei Cheng 等AAAI 2020 · 被引用 55 次
- Composing Photos Like a PhotographerChaoyi Hong, Shuaiyuan Du, Ke Xian, Hao Lu 等CVPR 2021
- Learning to Learn Cropping Models for Different Aspect Ratio RequirementsDebang Li, Junge Zhang, Kaiqi HuangCVPR 2020
相关 Paper
- Image Cropping with Spatial-aware Feature and Rank ConsistencyChao Wang, Li Niu, Bo Zhang, Liqing ZhangCVPR 2023
- Visualization Recommendation Through Visual Relation Learning and Visual Preference LearningDaomin Ji, Hui Luo, Zhifeng BaoICDE 2023 · 被引用 3 次
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou 等CVPR 2022 · 被引用 167 次
- Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image CompositionXiaoyu Liu, Ming Liu, Junyi Li, Shuai Liu 等ICCV 2023 · 被引用 7 次
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
