ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning
Ge Gao, Di Xiong, Zeke Xie, Jian Yang, Shuo Chen
Abstract
The unification of generative details and discriminative semantics presents a structural paradox in diffusion-based representation learning . Early approaches decouple semantics from generation, inevitably compromising representational completeness (i.e., information split ). While recent bridge-based methods achieve unification via a tightly coupled mapping, they suffer from information overload . This is because unconstrained reconstruction objectives incentivize the encoder to entangle high-frequency stochastic noise into the latent bottleneck. To solve this, we introduce asymmetric rectified contrastive diffusion autoencoder (ArcDAE), which rebuilds the diffusion bridge as a dynamic sifter . Through imposing a timestep-aware rectification constraint that orthogonalizes the semantic manifold from the stochastic noise space, ArcDAE compels the bottleneck to distill discriminative features while actively shedding high-frequency redundancy. Consequently, our approach eliminates the overload trap without reverting to decoupling. Extensive experiments validate the superiority of our FFHQ-trained ArcDAE, surpassing state-of-the-art methods by up to 6.4% in downstream semantics regression and 9.7% in reconstruction fidelity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55e61446-b4e1-45d5-a6f3-fcc018c6db77Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Diffusion Bridge AutoEncoders for Unsupervised Representation LearningYeongmin Kim, Kwanghyeon Lee, Minsang Park, Byeonghu Na et al.ICLR 2025
- Unified Latent Space for Understanding and Generation via Semantic Auto-encoderXiaojie Li, Yang Zhao, Ming Li, Yancheng Zhang et al.CVPR 2026
- Denoising Diffusion Autoencoders are Unified Self-supervised LearnersWeilai Xiang, Hongyu Yang, Di Huang, Yunhong WangICCV 2023 · 145 citations
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and GenerationXiang Wang, Zhifei Zhang, He Zhang, Zhe Lin et al.CVPR 2026 · 12 citations
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent SpaceJunyu Chen, Dongyun Zou, Wenkun He, Junsong Chen et al.ICCV 2025 · 3 citations
