LABridge: Text-Image Latent Alignment Framework via Mean-Conditioned OU Process
Huiyang Shao, Xin Xia, Yuxi Ren, Xing Wang, Xuefeng Xiao
摘要
Diffusion models have emerged as state-of-the-art in image synthesis.However, it often suffer from semantic instability and slow iterative denoising. We introduce Latent Alignment Framework (LABridge), a novel Text-Image Latent Alignment Framework via an Ornstein-Uhlenbeck (OU) Process, which explicitly preserves and aligns textual and visual semantics in an aligned latent space. LABridge employs a Text-Image Alignment Encoder (TIAE) to encode text prompts into structured priors that are directly aligned with image latents. Instead of a homogeneous Gaussian, Mean-Conditioned OU process smoothly interpolates between these text-conditioned priors and image latents, improving stability and reducing the number of denoising steps. Extensive experiments on standard text-to-image benchmarks show that LABridge achieves better text-image alignment metric and competitive FID scores compared to leading diffusion baselines. By unifying text and image representations through principled latent alignment, LABridge paves the way for more efficient, semantically consistent, and high-fidelity text to image generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image GenerationZijing Hu, Yunze Tong, Fengda Zhang, Junkun Yuan 等ICLR 2026 · 被引用 3 次
- Be Decisive: Noise-Induced Layouts for Multi-Subject GenerationOmer Dahary, Yehonathan Cohen, Or Patashnik, Kfir Aberman 等SIGGRAPH 2025 · 被引用 3 次
- DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination CapabilityRunhui Huang, Jianhua Han, Guansong Lu, Xiaodan Liang 等ICCV 2023 · 被引用 10 次
- SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image GenerationPaul Grimal, Michaël Soumm, Hervé Le Borgne, Olivier Ferret 等AAAI 2026 · 被引用 1 次
- Diffusion Adaptive Text Embedding for Text-to-Image Diffusion ModelsByeonghu Na, Minsang Park, Gyuwon Sim, Donghyeok Shin 等NeurIPS 2025 · 被引用 8 次
