Locality Preserving Markovian Transition for Instance Retrieval
Jifei Luo, Wenzheng Wu, Hantao Yao, Lu Yu, Changsheng Xu
Abstract
Diffusion-based re-ranking methods are effective in modeling the data manifolds through similarity propagation in affinity graphs. However, positive signals tend to diminish over several steps away from the source, reducing discriminative power beyond local regions. To address this issue, we introduce the Locality Preserving Markovian Transition (LPMT) framework, which employs a long-term thermodynamic transition process with multiple states for accurate manifold distance measurement. The proposed LPMT first integrates diffusion processes across separate graphs using Bidirectional Collaborative Diffusion (BCD) to establish strong similarity relationships. Afterwards, Locality State Embedding (LSE) encodes each instance into a distribution for enhanced local consistency. These distributions are interconnected via the Thermodynamic Markovian Transition (TMT) process, enabling efficient global retrieval while maintaining local effectiveness. Experimental results across diverse tasks confirm the effectiveness of LPMT for instance retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8bf3c92c-c6d2-44d8-86d0-a1350c49a961Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Cluster-Aware Similarity Diffusion for Instance RetrievalJifei Luo, Hantao Yao, Changsheng XuICML 2024 · 1 citation
- DiffTMR: Diffusion-based Hierarchical Alignment for Text-Molecule RetrievalChenxu Wang, Dong Zhou, Ting Liu, Jianghao Lin et al.ACM MM 2025
- GeoDiff: A Geometric Diffusion Model for Molecular Conformation GenerationMinkai Xu, Lantao Yu, Yang Song, Chence Shi et al.ICLR 2022 · 695 citations
- DiffusionRet: Generative Text-Video Retrieval with Diffusion ModelPeng Jin, Hao Li, Zesen Cheng, Kehan Li et al.ICCV 2023 · 95 citations
- Long Context Modeling with Ranked Memory-Augmented RetrievalGhadir Alselwi, Hao Xue, Shoaib Jameel, Basem Suleiman et al.ACL 2026 · 3 citations
