Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
Shilong Zhang, He Zhang, Zhifei Zhang, Chongjian GE, Shuchen Xue, Shaoteng Liu, Mengwei Ren, Soo Ye Kim, Yuqian Zhou, Qing Liu, Daniil Pakhomov, Kai Zhang
摘要
Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder's inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic-pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with 16 × 16 spatial downsampling). This design ensures that the latent space remains semantically rich and achieves state-of-the-art image reconstruction, while remaining compact enough for accurate generation. Leveraging this representation, we design a unified Text-to-Image (T2I) and image editing model. Benchmarking against various feature spaces, we demonstrate that our approach achieves state-of-the-art reconstruction, faster convergence, and substantial performance gains in both T2I and editing tasks, validating that representation encoders can be effectively adapted into robust generative components.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery AttacksXin Zhang, Zijin Yang, Kejiang Chen, Linfeng Ma 等ICML 2026 · 被引用 2 次
- RePack then Refine: Efficient Diffusion Transformers with Vision Foundation ModelsGuanfang Dong, Luke Schultz, Negar Hassanpour, Chao GaoICML 2026 · 被引用 1 次
- ForensicConcept: Transferable Forensic Concepts for AIGI DetectionMenyanshu Zhou, Ziyin Zhou, Ke Sun, Yunpeng Luo 等ICML 2026 · 被引用 1 次
- Concept-Guided Tokenization: Closing the Gap Between Reconstruction and GenerationYunqiao Yang, Haokun Lin, Guanzhong Wu, Ying WeiICML 2026
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and VerifyingChengyu Bai, Jintao Chen, Xiang Bai, Yilong Chen 等CVPR 2026 · 被引用 8 次
- Unified Latent Space for Understanding and Generation via Semantic Auto-encoderXiaojie Li, Yang Zhao, Ming Li, Yancheng Zhang 等CVPR 2026
- Boosting Generative Image Modeling via Joint Image-Feature SynthesisTheodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris 等NeurIPS 2025 · 被引用 47 次
- Decoder-Only LLMs are Better Controllers for Diffusion ModelsZiyi Dong, Yao Xiao, Pengxu Wei, Liang LinACM MM 2024 · 被引用 3 次
- LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image GenerationMushui Liu, Yuhang Ma, Zhen Yang, Jun Dan 等AAAI 2025 · 被引用 36 次
