Fully Functional Image Manipulation Using Scene Graphs in A Bounding-Box Free Way
Sitong Su, Lianli Gao, Junchen Zhu, Jie Shao, Jingkuan Song
Abstract
Recently, performing semantic editing of an image by modifying a scene graph has been proposed to support high-level image manipulation, and plays an important role for image generation. However, existing methods are all based on bounding boxes, and they suffer from the bounding box constraint. First, a bounding box often involves other instances (e.g, objects or environments) which do not need to be modified, but existing methods manipulate all the contents included in the bounding box. Secondly, prior methods fail to support adding instances when the bounding box of the target instance cannot be provided. To address the two issues above, we propose a novel bounding box free approach, which consists of two parts: a Local Bounding Box Free (Local-BBox-Free) Mask Generation and a Global Bounding Box Free (Global-BBox-Free) Instance Generation. The first part relieves the model of reliance on bounding boxes by generating the mask of the target instance to be manipulated without using the target instance bounding box. This enables our method to be the first to support fully functional image manipulation using scene graphs, including adding, removing, replacing and repositing instances. The second part is designed to synthesize the target instance directly from the generated mask and then paste it back to the inpainted original image using the generated mask, which preserves the unchanged part to the largest extent and precisely controls the target instance generation. Extensive experiments on Visual Genome and COCO-Stuff demonstrate that our model significantly surpasses the state-of-the-art both quantitatively and qualitatively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get abd51db2-329a-41f3-a742-b4fd12bd5835Cited by top-tier papers3
- Improving Scene Graph Generation with Superpixel-Based Interaction LearningJingyi Wang, Can Zhang, Jinfa Huang, Botao Ren et al.ACM MM 2023 · 8 citations
- ASSET: autoregressive semantic scene editing with transformers at high resolutionsDifan Liu, Sandesh Shetty, Tobias Hinz, Matthew Fisher et al.SIGGRAPH 2022 · 5 citations
- Generating Handwritten Mathematical Expressions From Symbol Graphs: An End-to-End PipelineYu Chen, Fei Gao, Yanguang Zhang, Maoying Qiao et al.CVPR 2024
Related papers
- Semantic Image Manipulation Using Scene GraphsHelisa Dhamo, Azade Farshad, Iro Laina, Nassir Navab et al.CVPR 2020
- Draw2Edit: Mask-Free Sketch-Guided Image ManipulationYiwen Xu, Ruoyu Guo, Maurice Pagnucco, Yang SongACM MM 2023 · 3 citations
- Hierarchical Image Generation via Transformer-Based Sequential Patch SelectionXiaogang Xu, Ning XuAAAI 2022 · 10 citations
- SketchEdit: Mask-Free Local Image Manipulation with Partial SketchesYu Zeng, Zhe Lin, Vishal M. PatelCVPR 2022 · 50 citations
- Object-Centric Image Generation from LayoutsTristan Sylvain, Pengchuan Zhang, Yoshua Bengio, R. Devon Hjelm et al.AAAI 2021 · 107 citations
