Spatial-Semantic Collaborative Cropping for User Generated Content
Yukun Su, Yiwen Cao, Jingliang Deng, Fengyun Rao, Qingyao Wu
摘要
A large amount of User Generated Content (UGC) is uploaded to the Internet daily and displayed to people worldwidely through the client side (e.g., mobile and PC). This requires the cropping algorithms to produce the aesthetic thumbnail within a specific aspect ratio on different devices. However, existing image cropping works mainly focus on landmark or landscape images, which fail to model the relations among the multi-objects with the complex background in UGC. Besides, previous methods merely consider the aesthetics of the cropped images while ignoring the content integrity, which is crucial for UGC cropping. In this paper, we propose a Spatial-Semantic Collaborative cropping network (S 2 CNet) for arbitrary user generated content accompanied by a new cropping benchmark. Specifically, we first mine the visual genes of the potential objects. Then, the suggested adaptive attention graph recasts this task as a procedure of information association over visual nodes. The underlying spatial and semantic relations are ultimately centralized to the crop candidate through differentiable message passing, which helps our network efficiently to preserve both the aesthetics and the content integrity. Extensive experiments on the proposed UGCrop5K and other public datasets demonstrate the superiority of our approach over state-of-the-art counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PhotoFramer: Multi-modal Image Composition InstructionZhiyuan You, Ke Wang, He Zhang, Xin Cai 等CVPR 2026 · 被引用 8 次
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and CroppingTianxiang Du, Hulingxiao He, Yuxin PengCVPR 2026 · 被引用 3 次
- ProCrop: Learning Aesthetic Image Cropping from Professional CompositionsKe Zhang, Tianyu Ding, Jiachen Jiang, Tianyi Chen 等AAAI 2026 · 被引用 3 次
- Towards Smart Point-and-Shoot PhotographyJiawan Li, Fei Zhou, Zhipeng Zhong, Jiongzhi Lin 等CVPR 2025
- Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and UnderstandingZhaoran Zhao, Peng Lu, Anran Zhang, Peipei Li 等CVPR 2025
它引用的顶会 Paper13
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding 等ICML 2020 · 被引用 1,910 次
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- Zero-Shot Video Object Segmentation via Attentive Graph Neural NetworksWenguan Wang, Xiankai Lu, Jianbing Shen, David J. Crandall 等ICCV 2019 · 被引用 294 次
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian 等CVPR 2022 · 被引用 132 次
- Image Cropping with Composition and Saliency Aware Aesthetic Score MapYi Tu, Li Niu, Weijie Zhao, Dawei Cheng 等AAAI 2020 · 被引用 55 次
相关 Paper
- Image Cropping with Spatial-aware Feature and Rank ConsistencyChao Wang, Li Niu, Bo Zhang, Liqing ZhangCVPR 2023
- Devil's on the Edges: Selective Quad Attention for Scene Graph GenerationDeunsol Jung, Sanghyun Kim, Won Hwa Kim, Minsu ChoCVPR 2023
- A Deep Learning based No-reference Quality Assessment Model for UGC VideosWei Sun, Xiongkuo Min, Wei Lu, Guangtao ZhaiACM MM 2022 · 被引用 239 次
- Find Beauty in the Rare: Contrastive Composition Feature Clustering for Nontrivial Cropping Box RegressionZhiyu Pan, Yinpeng Chen, Jiale Zhang, Hao Lu 等AAAI 2023 · 被引用 12 次
- Group Collaborative Learning for Co-Salient Object DetectionQi Fan, Deng-Ping Fan, Huazhu Fu, Chi-Keung Tang 等CVPR 2021
