Where's Waldo: Diffusion Features For Personalized Segmentation and Retrieval
Dvir Samuel, Rami Ben-Ari, Matan Levy, Nir Darshan, Gal Chechik
Abstract
Personalized retrieval and segmentation aim to locate specific instances within a dataset based on an input image and a short description of the reference instance. While supervised methods are effective, they require extensive labeled data for training. Recently, self-supervised foundation models have been introduced to these tasks showing comparable results to supervised methods. However, a significant flaw in these models is evident: they struggle to locate a desired instance when other instances within the same class are presented. In this paper, we explore text-to-image diffusion models for these tasks. Specifically, we propose a novel approach called PDM for Personalized Features Diffusion Matching, that leverages intermediate features of pre-trained text-to-image models for personalization tasks without any additional training. PDM demonstrates superior performance on popular retrieval and segmentation benchmarks, outperforming even supervised methods. We also highlight notable shortcomings in current instance and segmentation datasets and propose new benchmarks for these tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b278402-31a0-4673-aafb-cdb8c8d2f79aCited by top-tier papers10
- INSID3: Training-Free In-Context Segmentation with DINOv3Claudia Cuttano, Gabriele Trivigno, Christoph Reich, Daniel Cremers et al.CVPR 2026 · 13 citations
- Toward Early Quality Assessment of Text-to-Image Diffusion ModelsHuanlei Guo, Hongxin Wei, Bingyi JingCVPR 2026 · 2 citations
- Mind the Gap: Aligning Vision Foundation Models to Image Feature MatchingYuhan Liu, Jingwen Fu, Yang Wu, Kangyi Wu et al.ICCV 2025 · 2 citations
- Teaching VLMs to Localize Specific Objects from In-Context ExamplesSivan Doveh, Nimrod Shabtay, Eli Schwartz, Hilde Kuehne et al.ICCV 2025 · 2 citations
- Find your Needle: Small Object Image Retrieval via Multi-Object Attention OptimizationMichael Green, Matan Levy, Issar Tzachor, Dvir Samuel et al.NeurIPS 2025 · 1 citation
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Explore In-Context Segmentation via Latent Diffusion ModelsChaoyang Wang, Xiangtai Li, Henghui Ding, Lu Qi et al.AAAI 2025 · 17 citations
- DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image PersonalizationJisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin et al.CVPR 2024 · 21 citations
- Teleportraits: Training-Free People Insertion Into Any SceneJialu Gao, K. J. Joseph, Fernando De la TorreICCV 2025
- Conditional Latent Diffusion Models for Zero-Shot Instance SegmentationMaximilian Ulmer, Wout Boerdijk, Rudolph Triebel, Maximilian DurnerICCV 2025 · 1 citation
- Foreground-Background Separation through Concept Distillation from Generative Image Foundation ModelsMischa Dombrowski, Hadrien Reynaud, Matthew Baugh, Bernhard KainzICCV 2023 · 9 citations
