ID-Sim: An Identity-Focused Similarity Metric
Julia Chae, Nick Kolkin, Jui-Hsien Wang, Richard Zhang, Sara Beery, Cusuh Ham
Abstract
Humans have remarkable selective sensitivity to identities--they easily distinguish between highly-similar identities, even across significantly different contexts such as diverse viewpoints or lighting. Vision models have struggled to match this capability, and progress towards identity-focused tasks such as personalized image generation is slowed by a lack of identity-focused evaluation metrics. To help facilitate progress, we propose ID-Sim, a feed-forward metric designed to faithfully reflect human selective sensitivity. To build ID-Sim, we curate a high-quality training set of images spanning diverse real-world domains, augmented with generative synthetic data that provides controlled, fine-grained identity and contextual variations. We evaluate our metric on a new unified evaluation benchmark for assessing consistency with human annotations across identity-focused recognition, retrieval, and generative tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86f195ed-3a34-412d-9351-c850f5a2435fBuilds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Diffusion Models Through a Global Lens: Are They Culturally Inclusive?Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad et al.ACL 2025
- Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching ParadigmYufeng Cheng, wenxu wu, Shaojin Wu, Mengqi Huang et al.CVPR 2026
- Rethinking Deep Face RestorationYang Zhao, Yu-Chuan Su, Chun-Te Chu, Yandong Li et al.CVPR 2022 · 20 citations
- Illuminating Visual Identity in Universal Multimodal EmbeddingsJiawei Cao, Junyi Feng, Jiashen Hua, Ziheng Huang et al.CVPR 2026 · 1 citation
- Visual Personalization Turing TestRameen Abdal, James Burgess, Sergey Tulyakov, Kuan-Chieh Jackson WangCVPR 2026
