ASSET: autoregressive semantic scene editing with transformers at high resolutions
Difan Liu, Sandesh Shetty, Tobias Hinz, Matthew Fisher, Richard Zhang, Taesung Park, Evangelos Kalogerakis
摘要
We present ASSET, a neural architecture for automatically modifying an input high-resolution image according to a user's edits on its semantic segmentation map. Our architecture is based on a transformer with a novel attention mechanism. Our key idea is to sparsify the transformer's attention matrix at high resolutions, guided by dense attention extracted at lower image resolutions. While previous attention mechanisms are computationally too expensive for handling high-resolution images or are overly constrained within specific image regions hampering long-range interactions, our novel attention mechanism is both computationally efficient and effective. Our sparsified attention mechanism is able to capture long-range interactions and context, leading to synthesizing interesting phenomena in scenes, such as reflections of landscapes onto water or fora consistent with the rest of the landscape, that were not possible to generate reliably with previous convnets and transformer approaches. We present qualitative and quantitative results, along with user studies, demonstrating the effectiveness of our method. Our code and dataset are available at our project page: https://github.com/DifanLiu/ASSET
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TransLO: A Window-Based Masked Point Transformer Framework for Large-Scale LiDAR OdometryJiuming Liu, Guangming Wang, Chaokang Jiang, Zhe Liu 等AAAI 2023 · 被引用 56 次
- ClimateNeRF: Extreme Weather Synthesis in Neural Radiance FieldYuan Li, Zhi-Hao Lin, David A. Forsyth, Jia-Bin Huang 等ICCV 2023 · 被引用 44 次
- SIEDOB: Semantic Image Editing by Disentangling Object and BackgroundWuyang Luo, Su Yang, Xinjian Zhang, Weishan ZhangCVPR 2023
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- Taming Transformers for High-Resolution Image SynthesisPatrick Esser, Robin Rombach, Björn OmmerCVPR 2021
- HyperTransformer: A Textural and Spectral Feature Fusion Transformer for PansharpeningWele Gedara Chaminda Bandara, Vishal M. PatelCVPR 2022 · 被引用 175 次
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi 等CVPR 2024 · 被引用 137 次
- T-former: An Efficient Transformer for Image InpaintingYe Deng, Siqi Hui, Sanping Zhou, Deyu Meng 等ACM MM 2022 · 被引用 58 次
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi 等CVPR 2022 · 被引用 382 次
