Modeling Image Composition for Complex Scene Generation
Zuopeng Yang, Daqing Liu, Chaoyue Wang, Jie Yang, Dacheng Tao
摘要
We present a method that achieves state-of-the-art results on challenging (few-shot) layout-to-image generation tasks by accurately modeling textures, structures and relationships contained in a complex scene. After compressing RGB images into patch tokens, we propose the Transformer with Focal Attention (TwFA) for exploring dependencies of object-to-object, object-to-patch and patch-to-patch. Compared to existing CNN-based and Transformer-based generation models that entangled modeling on pixel-level&patch-level and object-level&patch-level respectively, the proposed focal attention predicts the current patch token by only focusing on its highly-related tokens that specified by the spatial layout, thereby achieving disambiguation during training. Furthermore, the proposed TwFA largely increases the data efficiency during training, therefore we propose the first few-shot complex scene generation strategy based on the well-trained TwFA. Comprehensive experiments show the superiority of our method, which significantly increases both quantitative metrics and qualitative visual realism with respect to state-of-the-art CNN-based and transformer-based methods. Code is available at https://github.com/JohnDreamer/TwFA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained DiffusionJinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu 等ICCV 2023 · 被引用 313 次
- Frido: Feature Pyramid Diffusion for Complex Scene Image SynthesisWan-Cyuan Fan, Yen-Chun Chen, Dongdong Chen, Yu Cheng 等AAAI 2023 · 被引用 118 次
- Directed Diffusion: Direct Control of Object Placement through Attention GuidanceWan-Duo Kurt Ma, Avisek Lahiri, John P. Lewis, Thomas Leung 等AAAI 2024 · 被引用 87 次
- Training-Free Structured Diffusion Guidance for Compositional Text-to-Image SynthesisWeixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani 等ICLR 2023 · 被引用 70 次
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsHanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan 等ICLR 2024 · 被引用 61 次
它引用的顶会 Paper17
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 被引用 429 次
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu 等ICCV 2021 · 被引用 427 次
- Few-shot Image Generation with Elastic Weight ConsolidationYijun Li, Richard Zhang, Jingwan Lu, Eli ShechtmanNeurIPS 2020 · 被引用 193 次
- Image Synthesis From Reconfigurable Layout and StyleWei Sun, Tianfu WuICCV 2019 · 被引用 160 次
相关 Paper
- Focus Your Attention when Few-Shot ClassificationHaoqing Wang, Shibo Jie, Zhihong DengNeurIPS 2023 · 被引用 16 次
- LayoutTransformer: Scene Layout Generation With Conceptual and Spatial DiversityCheng-Fu Yang, Wan-Cyuan Fan, Fu-En Yang, Yu-Chiang Frank WangCVPR 2021
- LayoutTransformer: Layout Generation and Completion with Self-attentionKamal Gupta, Justin Lazarow, Alessandro Achille, Larry Davis 等ICCV 2021 · 被引用 184 次
- Learned Spatial Representations for Few-shot Talking-Head SynthesisMoustafa Meshry, Saksham Suri, Larry S. Davis, Abhinav ShrivastavaICCV 2021 · 被引用 51 次
- MaskSketch: Unpaired Structure-guided Masked Image GenerationDina Bashkirova, José Lezama, Kihyuk Sohn, Kate Saenko 等CVPR 2023
