Generic Perceptual Loss for Modeling Structured Output Dependencies
Yifan Liu, Hao Chen, Yu Chen, Wei Yin, Chunhua Shen
Abstract
The perceptual loss has been widely used as an effective loss term in image synthesis tasks including image superresolution [16] , and style transfer [14] . It was believed that the success lies in the high-level perceptual feature representations extracted from CNNs pretrained with a large set of images. Here we reveal that, what matters is the network structure instead of the trained weights. Without any learning, the structure of a deep network is sufficient to capture the dependencies between multiple levels of variable statistics using multiple layers of CNNs. This insight removes the requirements of pre-training and a particular network structure (commonly, VGG) that are previously assumed for the perceptual loss, thus enabling a significantly wider range of applications. To this end, we demonstrate that a randomly-weighted deep CNN can be used to model the structured dependencies of outputs. On a few dense perpixel prediction tasks such as semantic segmentation, depth estimation and instance segmentation, we show improved results of using the extended randomized perceptual loss, compared to the baselines using pixel-wise loss alone. We hope that this simple, extended perceptual loss may serve as a generic structured-output loss that is applicable to most structured output learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5674c985-d980-402d-89bc-1cc3b2dd46f2Cited by top-tier papers1
Ask how each one uses itBuilds on3
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- BlendMask: Top-Down Meets Bottom-Up for Instance SegmentationHao Chen, Kunyang Sun, Zhi Tian, Chunhua Shen et al.CVPR 2020
Related papers
- Understanding and Simplifying Perceptual DistancesDan Amir, Yair WeissCVPR 2021
- You Only Need Adversarial Supervision for Semantic Image SynthesisEdgar Schönfeld, Vadim Sushko, Dan Zhang, Juergen Gall et al.ICLR 2021 · 219 citations
- SROBB: Targeted Perceptual Loss for Single Image Super-ResolutionMohammad Saeed Rad, Behzad Bozorgtabar, Urs-Viktor Marti, Max Basler et al.ICCV 2019 · 147 citations
- Towards Interpretable Face RecognitionBangjie Yin, Luan Tran, Haoxiang Li, Xiaohui Shen et al.ICCV 2019 · 92 citations
- Context Prior for Scene SegmentationChangqian Yu, Jingbo Wang, Changxin Gao, Gang Yu et al.CVPR 2020
