TopNet: Transformer-Based Object Placement Network for Image Compositing
Sijie Zhu, Zhe Lin, Scott Cohen, Jason Kuen, Zhifei Zhang, Chen Chen
Abstract
We investigate the problem of automatically placing an object into a background image for image compositing. Given a background image and a segmented object, the goal is to train a model to predict plausible placements (location and scale) of the object for compositing. The quality of the composite image highly depends on the predicted location/scale. Existing works either generate candidate bounding boxes or apply sliding-window search using global representations from background and object images, which fail to model local information in background images. However, local clues in background images are important to determine the compatibility of placing the objects with certain locations/scales. In this paper, we propose to learn the correlation between object features and all local background features with a transformer module so that detailed information can be provided on all possible location/scale configurations. A sparse contrastive loss is further proposed to train our model with sparse supervision. Our new formulation generates a 3D heatmap indicating the plausibility of all location/scale combinations in one network forward pass, which is > 10× faster than the previous slidingwindow method. It also supports interactive search when users provide a pre-defined location or scale. The proposed method can be trained with explicit annotation or in a self-supervised manner using an off-the-shelf inpainting model, and it outperforms state-of-the-art methods significantly. User study shows that the trained model generalizes well to real-world images with diverse challenging scenes and object categories.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Placeit3d: Language-Guided Object Placement in Real 3D ScenesAhmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi et al.ICCV 2025 · 11 citations
- SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout ControlJaskirat Singh, Jianming Zhang, Qing Liu, Cameron Smith et al.CVPR 2024 · 7 citations
- MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular DetectionRishubh Parihar, Srinjay Sarkar, Sarthak Vora, Jogendra Nath Kundu et al.CVPR 2025
- BOOTPLACE: Bootstrapped Object Placement with Detection TransformersHang Zhou, Xinxin Zuo, Rui Ma, Li ChengCVPR 2025
- CareCom: Generative Image Composition with Calibrated Reference FeaturesJiaxuan Chen, Bo Zhang, Qingdong He, Jinlong Peng et al.AAAI 2026
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With TransformersSixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu et al.CVPR 2021
- CenterMask: Real-Time Anchor-Free Instance SegmentationYoungwan Lee, Jongyoul ParkCVPR 2020
Related papers
- HEAP: Unsupervised Object Discovery and Localization with Contrastive GroupingXin Zhang, Jinheng Xie, Yuan Yuan, Michael Bi Mi et al.AAAI 2024 · 11 citations
- Contrastive Attention Maps for Self-supervised Co-localizationMinsong Ki, Youngjung Uh, Junsuk Choe, Hyeran ByunICCV 2021 · 11 citations
- Proposal-Contrastive Pretraining for Object Detection from Fewer DataQuentin Bouniot, Romaric Audigier, Angélique Loesch, Amaury HabrardICLR 2023
- Interactive Object Placement with Reinforcement LearningShengping Zhang, Quanling Meng, Qinglin Liu, Liqiang Nie et al.ICML 2023 · 9 citations
- Distilling Localization for Self-Supervised Representation LearningNanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinAAAI 2021 · 59 citations
