Neural Congealing: Aligning Images to a Joint Semantic Atlas
Dolev Ofri-Amar, Michal Geyer, Yoni Kasten, Tali Dekel
Abstract
Figure 1 . Given a set of input images, our method automatically detects and jointly aligns semantically-common content across the images. This is achieved through a test-time training approach that estimates a unified 2D atlas that represents the common semantic content, and dense mappings from the joint atlas to each of the input images. Our atlas and mappings are optimized per input set in a self-supervised manner by leveraging a pre-trained DINO-ViT model. Our method can be applied to diverse image sets, without requiring any additional training data, and allows us to automatically propagate an edit applied to a single image across the entire set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19940444-c33f-4eab-8254-f010b81dbdaaCited by top-tier papers21
- Emergent Correspondence from Image DiffusionLuming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo et al.NeurIPS 2023 · 555 citations
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera et al.NeurIPS 2023 · 371 citations
- Diffusion Hyperfeatures: Searching Through Time and Space for Semantic CorrespondenceGrace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski et al.NeurIPS 2023 · 261 citations
- Cross-Image Attention for Zero-Shot Appearance TransferYuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor et al.SIGGRAPH 2024 · 72 citations
- ASIC: Aligning Sparse in-the-wild Image CollectionsKamal Gupta, Varun Jampani, Carlos Esteves, Abhinav Shrivastava et al.ICCV 2023 · 30 citations
Builds on18
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Semantic Segmentation by Distilling Feature CorrespondencesMark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely et al.ICLR 2022 · 317 citations
- Generalized Category DiscoverySagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanCVPR 2022 · 194 citations
- Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutYangtao Wang, Xi Shen, Shell Xu Hu, Yuan Yuan et al.CVPR 2022 · 143 citations
- CATs: Cost Aggregation Transformers for Visual CorrespondenceSeokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee et al.NeurIPS 2021 · 133 citations
Related papers
- Finding Distributed Object-Centric Properties in Self-Supervised TransformersSamyak Rawlekar, Amitabh Swain, Yujun Cai, Yiwei Wang et al.CVPR 2026 · 1 citation
- Splicing ViT Features for Semantic Appearance TransferNarek Tumanyan, Omer Bar-Tal, Shai Bagon, Tali DekelCVPR 2022 · 128 citations
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 52 citations
- Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object DetectionMarc-Antoine Lavoie, Anas Mahmoud, Steven L. WaslanderCVPR 2025
- Revisiting Change Captioning from Self-supervised Global-Part AlignmentFeixiao Lv, Rui Wang, Lihua JingAAAI 2025 · 1 citation
