Interactive Object Placement with Reinforcement Learning
Shengping Zhang, Quanling Meng, Qinglin Liu, Liqiang Nie, Bineng Zhong, Xiaopeng Fan, Rongrong Ji
摘要
Object placement aims to insert a foreground object into a background image with a suitable location and size to create a natural composition. To predict a diverse distribution of placements, existing methods usually establish a oneto-one mapping from random vectors to the placements. However, these random vectors are not interpretable, which prevents users from interacting with the object placement process. To address this problem, we propose an Interactive Object Placement method with Reinforcement Learning, dubbed IOPRE, to make sequential decisions for producing a reasonable placement given an initial location and size of the foreground. We first design a novel action space to flexibly and stably adjust the location and size of the foreground while preserving its aspect ratio. Then, we propose a multi-factor state representation learning method, which integrates composition image features and sinusoidal positional embeddings of the foreground to make decisions for selecting actions. Finally, we design a hybrid reward function that combines placement assessment and the number of steps to ensure that the agent learns to place objects in the most visually pleasing and semantically appropriate location. Experimental results on the OPA dataset demonstrate that the proposed method achieves state-of-the-art performance in terms of plausibility and diversity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- InstaBoost: Boosting Instance Segmentation via Probability Map Guided Copy-PastingHaoshu Fang, Jianhua Sun, Runzhong Wang, Minghao Gou 等ICCV 2019 · 被引用 236 次
- Temporal Complementarity-Guided Reinforcement Learning for Image-to-Video Person Re-IdentificationWei Wu, Jiawei Liu, Kecheng Zheng, Qibin Sun 等CVPR 2022 · 被引用 17 次
- MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image SynthesisShuchen Weng, Wenbo Li, Dawei Li, Hongxia Jin 等CVPR 2020
相关 Paper
- TopNet: Transformer-Based Object Placement Network for Image CompositingSijie Zhu, Zhe Lin, Scott Cohen, Jason Kuen 等CVPR 2023
- BOOTPLACE: Bootstrapped Object Placement with Detection TransformersHang Zhou, Xinxin Zuo, Rui Ma, Li ChengCVPR 2025
- HOComp: Interaction-Aware Human-Object CompositionDong Liang, Jinyuan Jia, Yuhao Liu, Rynson W. H. LauNeurIPS 2025 · 被引用 1 次
- Self-Play Reinforcement Learning for Fast Image RetargetingNobukatsu Kajiura, Satoshi Kosugi, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 被引用 21 次
- ORIDa: Object-centric Real-world Image Composition DatasetJinwoo Kim, Sangmin Han, Jinho Jeong, Jiwoo Choi 等CVPR 2025
