Learning Generative Vision Transformer with Energy-Based Latent Space for Saliency Prediction
Jing Zhang, Jianwen Xie, Nick Barnes, Ping Li
摘要
Vision transformer networks have shown superiority in many computer vision tasks. In this paper, we take a step further by proposing a novel generative vision transformer with latent variables following an informative energy-based prior for salient object detection. Both the vision transformer network and the energy-based prior model are jointly trained via Markov chain Monte Carlo-based maximum likelihood estimation, in which the sampling from the intractable posterior and prior distributions of the latent variables are performed by Langevin dynamics. Further, with the generative vision transformer, we can easily obtain a pixel-wise uncertainty map from an image, which indicates the model confidence in predicting saliency from the image. Different from the existing generative models which define the prior distribution of the latent variables as a simple isotropic Gaussian distribution, our model uses an energy-based informative prior which can be more expressive to capture the latent space of the data. We apply the proposed framework to both RGB and RGB-D salient object detection tasks. Extensive experimental results show that our framework can achieve not only accurate saliency predictions but also meaningful uncertainty maps that are consistent with the human perception. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- AVSegFormer: Audio-Visual Segmentation with TransformerShengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang 等AAAI 2024 · 被引用 96 次
- Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationShilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen 等AAAI 2024 · 被引用 67 次
- CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video SegmentationKexin Li, Zongxin Yang, Lei Chen, Yi Yang 等ACM MM 2023 · 被引用 58 次
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong 等ICCV 2023 · 被引用 57 次
- Joint Semantic Mining for Weakly Supervised RGB-D Salient Object DetectionJingjing Li, Wei Ji, Qi Bi, Cheng Yan 等NeurIPS 2021 · 被引用 56 次
它引用的顶会 Paper29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang 等ICCV 2021 · 被引用 1,062 次
相关 Paper
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao 等ICCV 2021 · 被引用 473 次
- UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational AutoencodersJing Zhang, Deng-Ping Fan, Yuchao Dai, Saeed Anwar 等CVPR 2020
- Energy-Based Generative Cooperative Saliency PredictionJing Zhang, Jianwen Xie, Zilong Zheng, Nick BarnesAAAI 2022 · 被引用 13 次
- TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding NetworkZhengyi Liu, Yuan Wang, Zhengzheng Tu, Yun Xiao 等ACM MM 2021 · 被引用 175 次
- Geometry-Free View Synthesis: Transformers and no 3D PriorsRobin Rombach, Patrick Esser, Björn OmmerICCV 2021 · 被引用 115 次
