Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
Shilong Zhang, He Zhang, Zhifei Zhang, Chongjian GE, Shuchen Xue, Shaoteng Liu, Mengwei Ren, Soo Ye Kim, Yuqian Zhou, Qing Liu, Daniil Pakhomov, Kai Zhang
Abstract
Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder's inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic-pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with 16 × 16 spatial downsampling). This design ensures that the latent space remains semantically rich and achieves state-of-the-art image reconstruction, while remaining compact enough for accurate generation. Leveraging this representation, we design a unified Text-to-Image (T2I) and image editing model. Benchmarking against various feature spaces, we demonstrate that our approach achieves state-of-the-art reconstruction, faster convergence, and substantial performance gains in both T2I and editing tasks, validating that representation encoders can be effectively adapted into robust generative components.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery AttacksXin Zhang, Zijin Yang, Kejiang Chen, Linfeng Ma et al.ICML 2026 · 2 citations
- RePack then Refine: Efficient Diffusion Transformers with Vision Foundation ModelsGuanfang Dong, Luke Schultz, Negar Hassanpour, Chao GaoICML 2026 · 1 citation
- ForensicConcept: Transferable Forensic Concepts for AIGI DetectionMenyanshu Zhou, Ziyin Zhou, Ke Sun, Yunpeng Luo et al.ICML 2026 · 1 citation
- Concept-Guided Tokenization: Closing the Gap Between Reconstruction and GenerationYunqiao Yang, Haokun Lin, Guanzhong Wu, Ying WeiICML 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and VerifyingChengyu Bai, Jintao Chen, Xiang Bai, Yilong Chen et al.CVPR 2026 · 8 citations
- Unified Latent Space for Understanding and Generation via Semantic Auto-encoderXiaojie Li, Yang Zhao, Ming Li, Yancheng Zhang et al.CVPR 2026
- Boosting Generative Image Modeling via Joint Image-Feature SynthesisTheodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris et al.NeurIPS 2025 · 47 citations
- Decoder-Only LLMs are Better Controllers for Diffusion ModelsZiyi Dong, Yao Xiao, Pengxu Wei, Liang LinACM MM 2024 · 3 citations
- LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image GenerationMushui Liu, Yuhang Ma, Zhen Yang, Jun Dan et al.AAAI 2025 · 36 citations
