ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting
Yizhe Tang, Zhimin Sun, Yuzhen Du, Ran Yi, Guangben Lu, Teng Hu, Luying Li, Lizhuang Ma, Fangyuan Zou
摘要
Image inpainting aims to fill the missing region of an image. Recently, there has been a surge of interest in foregroundconditioned background inpainting, a sub-task that fills the background of an image while the foreground subject and associated text prompt are provided. Existing background inpainting methods typically strictly preserve the subject's original position from the source image, resulting in inconsistencies between the subject and the generated background. To address this challenge, we propose a new task, the "Text-Guided Subject-Position Variable Background Inpainting", which aims to dynamically adjust the subject position to achieve a harmonious relationship between the subject and the inpainted background, and propose the Adaptive Transformation Agent (A T A) for this task. Firstly, we design a PosAgent Block that adaptively predicts an appropriate displacement based on given features to achieve variable subject-position. Secondly, we design the Reverse Displacement Transform (RDT) module, which arranges multiple PosAgent blocks in a reverse structure, to transform hierarchical feature maps from deep to shallow based on semantic information. Thirdly, we equip A T A with a Position Switch Embedding to control whether the subject's position in the generated image is adaptively predicted or fixed. Extensive comparative experiments validate the effectiveness of our A T A approach, which not only demonstrates superior inpainting capabilities in subject-position variable inpainting, but also ensures good performance on subjectposition fixed inpainting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Social Debiasing for Fair Multi-Modal LLMsHarry Cheng, Yangyang Guo, Qing Guo, Ming-Hsuan Yang 等ICCV 2025 · 被引用 1 次
- Pinco: Position-Induced Consistent Adapter for Diffusion Transformer in Foreground-Conditioned InpaintingGuangben Lu, Yuzhen Du, Yizhe Tang, Zhimin Sun 等ICCV 2025
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- InOut: Diverse Image Outpainting via GAN InversionYen-Chi Cheng, Chieh Hubert Lin, Hsin-Ying Lee, Jian Ren 等CVPR 2022 · 被引用 72 次
- Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image GeneratorChaehun Shin, Jooyoung Choi, Heeseung Kim, Sungroh YoonCVPR 2025
- Text-Guided Neural Image InpaintingLisai Zhang, Qingcai Chen, Baotian Hu, Shuoran JiangACM MM 2020 · 被引用 53 次
- DAFT-GAN: Dual Affine Transformation Generative Adversarial Network for Text-Guided Image InpaintingJihoon Lee, Yunhong Min, Hwidong Kim, Sangtae AhnACM MM 2024 · 被引用 3 次
- Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image GenerationTianyidan Xie, Rui Ma, Qian Wang, Xiaoqian Ye 等AAAI 2025 · 被引用 8 次
