Interactive Object Placement with Reinforcement Learning
Shengping Zhang, Quanling Meng, Qinglin Liu, Liqiang Nie, Bineng Zhong, Xiaopeng Fan, Rongrong Ji
Abstract
Object placement aims to insert a foreground object into a background image with a suitable location and size to create a natural composition. To predict a diverse distribution of placements, existing methods usually establish a oneto-one mapping from random vectors to the placements. However, these random vectors are not interpretable, which prevents users from interacting with the object placement process. To address this problem, we propose an Interactive Object Placement method with Reinforcement Learning, dubbed IOPRE, to make sequential decisions for producing a reasonable placement given an initial location and size of the foreground. We first design a novel action space to flexibly and stably adjust the location and size of the foreground while preserving its aspect ratio. Then, we propose a multi-factor state representation learning method, which integrates composition image features and sinusoidal positional embeddings of the foreground to make decisions for selecting actions. Finally, we design a hybrid reward function that combines placement assessment and the number of steps to ensure that the agent learns to place objects in the most visually pleasing and semantically appropriate location. Experimental results on the OPA dataset demonstrate that the proposed method achieves state-of-the-art performance in terms of plausibility and diversity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e46099eb-a601-44f6-8e5d-91c89d4c943cCited by top-tier papers1
Ask how each one uses itBuilds on4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-PastingHaoshu Fang, Jianhua Sun, Runzhong Wang, Minghao Gou et al.ICCV 2019 · 236 citations
- Temporal Complementarity-Guided Reinforcement Learning for Image-to-Video Person Re-IdentificationWei Wu, Jiawei Liu, Kecheng Zheng, Qibin Sun et al.CVPR 2022 · 17 citations
- MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image SynthesisShuchen Weng, Wenbo Li, Dawei Li, Hongxia Jin et al.CVPR 2020
Related papers
- TopNet: Transformer-Based Object Placement Network for Image CompositingSijie Zhu, Zhe Lin, Scott Cohen, Jason Kuen et al.CVPR 2023
- BOOTPLACE: Bootstrapped Object Placement with Detection TransformersHang Zhou, Xinxin Zuo, Rui Ma, Li ChengCVPR 2025
- HOComp: Interaction-Aware Human-Object CompositionDong Liang, Jinyuan Jia, Yuhao Liu, Rynson W. H. LauNeurIPS 2025 · 1 citation
- Self-Play Reinforcement Learning for Fast Image RetargetingNobukatsu Kajiura, Satoshi Kosugi, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 21 citations
- ORIDa: Object-centric Real-world Image Composition DatasetJinwoo Kim, Sangmin Han, Jinho Jeong, Jiwoo Choi et al.CVPR 2025
