Align Images Before You Generate
Shihua Zhang, Qiuhong Shen, Xinchao Wang
Abstract
respondences from the diffusion model's intermediate features, and an aligned area aggregator that integrates messages from only matching regions to avoid ambiguous information interactions. Given the native correspondences as guidance, CorrAdapter can enhance spatiotemporal consistency without any auxiliary inputs, and remains trainingfree and baseline-agnostic, which enables it to generalize seamlessly to various generation tasks. Additionally, we provide an optional training scheme to explore furtherimproved possibilities. Experiments on both static multiview generation and dynamic video generation show that CorrAdapter consistently improves spatiotemporal consistency and perceptual quality over strong baselines, offering a simple yet versatile drop-in approach to geometrically faithful multi-image diffusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on32
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
Related papers
- WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single ImageJiwoo Park, Tae Eun Choi, Youngjun Jun, Seong Jae HwangICCV 2025
- RefDrop: Controllable Consistency in Image or Video Generation via Reference Feature GuidanceJiaojiao Fan, Haotian Xue, Qinsheng Zhang, Yongxin ChenNeurIPS 2024 · 7 citations
- Correspondence-Attention Alignment for Multi-View Diffusion ModelsMinkyung Kwon, Jinhyeok Choi, Jiho Park, Seonghu Jeon et al.CVPR 2026
- ViewFusion: Towards Multi-View Consistency via Interpolated DenoisingXianghui Yang, Yan Zuo, Sameera Ramasinghe, Loris Bazzani et al.CVPR 2024 · 5 citations
- ConsistNet: Enforcing 3D Consistency for Multi-View Images DiffusionJiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji et al.CVPR 2024 · 25 citations
