ICML2026

Enhanced Latent-Space Adversarial Training for Super-Resolution

Liangbin Xie, Zheyuan Li, Fanghua Yu, Xinqi Lin, Jun-hao Zhuang, Jinfan Hu, Jinjin Gu, Jiantao Zhou, Chao Dong

摘要

Real-world super-resolution (SR) at large upscaling factors (i.e., ≥ 4×) remains difficult due to complex real-image degradations. HYPIR, a leading diffusion-based restoration model, performs strongly on many inputs, yet for a non-trivial portion of more challenging cases a single forward step does not fully recover fine-grained details. A naive two-stage cascade improves visual quality, but introduces over-saturation, weak texture details, and high inference latency. To address these issues, this paper proposes HYPIR++. It removes the degradation removal encoder and noise augmentation modules to better preserve fidelity cues from low-quality inputs. Equipped with an optimized latent ConvNeXt and a latent patch discriminator, HYPIR++ supports latentspace adversarial learning for clearer details and more stable local structures. It further shortens text sequences and replaces full attention with sparse neighbor attention, enabling efficient highresolution inference without tiling. Experiments show that HYPIR++ improves perceptual quality and runs 1.71× faster than HYPIR on large-factor real-world SR