TransView: Inside, Outside, and Across the Cropping View Boundaries
Zhiyu Pan, Zhiguo Cao, Kewei Wang, Hao Lu, Weicai Zhong
Abstract
We show that relation modeling between visual elements matters in cropping view recommendation. Cropping view recommendation addresses the problem of image recomposition conditioned on the composition quality and the ranking of views (cropped sub-regions). This task is challenging because the visual difference is subtle when a visual element is reserved or removed. Existing methods represent visual elements by extracting region-based convolutional features inside and outside the cropping view boundaries, without probing a fundamental question: why some visual elements are of interest or of discard? In this work, we observe that the relation between different visual elements significantly affects their relative positions to the desired cropping view, and such relation can be characterized by the attraction inside/outside the cropping view boundaries and the repulsion across the boundaries. By instantiating a transformer-based solution that represents visual elements as visual words and that models the dependencies between visual words, we report not only state-of-the-art performance on public benchmarks, but also interesting visualizations that depict the attraction and repulsion between visual elements, which may shed light on what makes for effective cropping view recommendation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext feb2ba16-9372-41c7-85e4-05d46615004dCited by top-tier papers9
- Rethinking Image Cropping: Exploring Diverse Compositions from Global ViewsGengyun Jia, Huaibo Huang, Chaoyou Fu, Ran HeCVPR 2022 · 19 citations
- Learning Second-Order Attentive Context for Efficient Correspondence PruningXinyi Ye, Weiyue Zhao, Hao Lu, Zhiguo CaoAAAI 2023 · 12 citations
- Shoot360: Normal View Video Creation from City Panorama FootageAnyi Rao, Linning Xu, Dahua LinSIGGRAPH 2022 · 5 citations
- ProCrop: Learning Aesthetic Image Cropping from Professional CompositionsKe Zhang, Tianyu Ding, Jiachen Jiang, Tianyi Chen et al.AAAI 2026 · 3 citations
- Exposure Completing for Temporally Consistent Neural High Dynamic Range Video RenderingJiahao Cui, Wei Jiang, Zhan Peng, Zhiyu Pan et al.ACM MM 2024 · 3 citations
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Image Cropping with Composition and Saliency Aware Aesthetic Score MapYi Tu, Li Niu, Weijie Zhao, Dawei Cheng et al.AAAI 2020 · 55 citations
- Composing Photos Like a PhotographerChaoyi Hong, Shuaiyuan Du, Ke Xian, Hao Lu et al.CVPR 2021
- Learning to Learn Cropping Models for Different Aspect Ratio RequirementsDebang Li, Junge Zhang, Kaiqi HuangCVPR 2020
Related papers
- Image Cropping with Spatial-aware Feature and Rank ConsistencyChao Wang, Li Niu, Bo Zhang, Liqing ZhangCVPR 2023
- Visualization Recommendation Through Visual Relation Learning and Visual Preference LearningDaomin Ji, Hui Luo, Zhifeng BaoICDE 2023 · 3 citations
- TransVPR: Transformer-Based Place Recognition with Multi-Level Attention AggregationRuotong Wang, Yanqing Shen, Weiliang Zuo, Sanping Zhou et al.CVPR 2022 · 167 citations
- Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image CompositionXiaoyu Liu, Ming Liu, Junyi Li, Shuai Liu et al.ICCV 2023 · 7 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
