InstaFormer: Instance-Aware Image-to-Image Translation with Transformer
Soohyun Kim, Jongbeom Baek, Jihye Park, Gyeongnyeon Kim, Seungryong Kim
摘要
We present a novel Transformer-based network architecture for instance-aware image-to-image translation, dubbed InstaFormer, to effectively integrate global- and instance-level information. By considering extracted content featuresfrom an image as tokens, our networks discover global consensus of content features by considering context information through a self-attention module in Transformers. By augmenting such tokens with an instance-level feature extracted from the content feature with respect to bounding box information, our framework is capable of learning an interaction between object instances and the global image, thus boosting the instance-awareness. We replace layer normalization (LayerNorm) in standard Transformers with adaptive instance normalization (AdaIN) to enable a multi-modal translation with style codes. In addition, to improve the instance-awareness and translation quality at object regions, we present an instance-level content contrastive loss defined between input and translated image. We conduct experiments to demonstrate the effectiveness of our InstaFormer over the latest methods and provide extensive ablation studies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SHUNIT: Style Harmonization for Unpaired Image-to-Image TranslationSeokbeom Song, Suhyeon Lee, Hongje Seong, Kyoungwon Min 等AAAI 2023 · 被引用 11 次
- PTUS: Photo-Realistic Talking Upper-Body Synthesis via 3D-Aware Motion Decomposition WarpingLuoyang Lin, Zutao Jiang, Xiaodan Liang, Liqian Ma 等AAAI 2024 · 被引用 3 次
- Bridging Day and Night: Target-Class Hallucination Suppression in Unpaired Image TranslationShuwei Li, Lei Tan, Robby T. TanAAAI 2026 · 被引用 2 次
- DMGINE: Day-Memory Guided Nighttime Image Enhancement for Dynamic Traffic ScenesRuizhou Liu, Zhe Wu, Zimo Liu, Qingfang Zheng 等AAAI 2026
- LANIT: Language-Driven Image-to-Image Translation for Unlabeled DataJihye Park, Sunwoo Kim, Soohyun Kim, Seokju Cho 等CVPR 2023
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- Adaptive Convolutions for Structure-Aware Style TransferPrashanth Chandran, Gaspard Zoss, Paulo F. U. Gotardo, Markus Gross 等CVPR 2021
- Translate the Facial Regions You Like Using Self-Adaptive Region TranslationWenshuang Liu, Wenting Chen, Zhanjia Yang, Linlin ShenAAAI 2021 · 被引用 9 次
- U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image TranslationJunho Kim, Minjae Kim, Hyeonwoo Kang, Kwanghee LeeICLR 2020 · 被引用 632 次
- Cross-Granularity Learning for Multi-Domain Image-to-Image TranslationHuiyuan Fu, Ting Yu, Xin Wang, Huadong MaACM MM 2020 · 被引用 4 次
- AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style TransferSonghua Liu, Tianwei Lin, Dongliang He, Fu Li 等ICCV 2021 · 被引用 421 次
