On the rankability of visual embeddings
Ankit Sonthalia, Arnas Uselis, Seong Joon Oh
Abstract
We study whether visual embedding models capture continuous, ordinal attributes along linear directions, which we term rank axes. We define a model as rankable for an attribute if projecting embeddings onto such an axis preserves the attribute's order. Across 7 popular encoders and 9 datasets with attributes like age, crowd count, head pose, aesthetics, and recency, we find that many embeddings are inherently rankable. Surprisingly, a small number of samples, or even just two extreme examples, often suffice to recover meaningful rank axes, without full-scale supervision. These findings open up new use cases for image ranking in vector databases and motivate further study into the structure and learning of rankable embeddings. Our code is available at https://github.com/aktsonthalia/rankable- vision-embeddings.
We examine two questions: (1) Are visual embeddings rankable? (2) How easily can we recover the rank axis for a given attribute? 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
To address (1), we evaluate 7 modern visual encoders, from ResNet to CLIP, across 9 datasets with 7 attributes: age, crowd count, 3 head pose angles (pitch, roll, yaw), image aesthetics, and recency. We find that many embedding spaces are indeed rankable (Section 3).
To address (2), we estimate the rank axis v A with minimal supervision. The structure of the embedding space makes full-dataset regression unnecessary. In many cases, a handful of annotated samples and, in some cases, a pair of samples x l (low) and x h (high) already recover non-trivial ranking performance. For the latter case, we define the rank axis as
This opens up the possibility for fast ordering of new images by arbitrary attributes. For example, a photo app lets users sort selfies by age appearance. It uses CLIP embeddings and two reference images: one of a child and one of an elderly person. The app computes v age without training. Users scroll from youngest-looking to oldest-looking faces in their album (Section 4).
Contributions:
- We define and motivate rankability as a property of visual embeddings, distinct from retrieval. 2. We study rankability across modern encoders and real-valued attributes; results show that current embeddings are rankable. 3. We show that rank axes can sometimes be recovered using only two or a handful of labelled samples.
2 Related work Embeddings for retrieval. Visual encoders are commonly used to index images in vector databases, enabling nearest neighbour search for retrieval tasks [42,52,33,57]. This setup, known as deep metric learning [5,6,42], predates vision-language models like CLIP [52]. CLIP and related models shifted the focus to cross-modal similarity modelling, where vision and language share a joint embedding space used for classification [52], retrieval [75,2], and captioning [39,26]. While the majority of work in visual encoders is devoted to the understanding of the local similarity structure, we study how visual embeddings support global ranking instead of just local retrieval.
Prior work has explored ways to improve the geometry of the embedding space. Order embeddings and hyperbolic representations have been used to model hierarchies [65,21,8,51]. Training disentangled representation [71] is considered critical for compositionality, where attributes are assigned to certain linear subspaces [60,3]. Others have defined concepts like uniformity and separability of the representations [70]. In this work, we focus on the analysis of a wide range of visual encoders, rather than introducing recipes for improvements.
A large body of work has examined the geometry and structure of CLIP's learned embedding space. CLIP and its derivatives have been studied extensively [4,53,32,77]. Several works have reported modality gaps between vision and language embeddings [12]. Some studies point to the absence of certain structures and capabilities in CLIP representations: attribute-object bindings [31,80,25], or the association of attributes to corresponding instances. Others argue that much information is already present in CLIP representations, including parts-of-speech and linguistic structure [44], attribute-object bindings [24], and compositional attributes [62,63]. The platonic representation hypothesis further suggests that models converge to similar internal structures [19]. In this work, we analyse the embedding geometry and structure for modern visual embeddings from the novel perspective of rankability.
Linearly probing an embedding. Linear probing is a fast and widely used method to test for the presence of concepts in visual embeddings [23,18,64]. It measures the accuracy of a linear classifier trained on intermediate-layer features, effectively testing whether a hyperplane can separate embeddings containing a concept from those that do not. This technique has been used to study the geometry of CLIP's embeddings [30] and to probe for specific information such as attribute-objec
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c89bfa96-982e-4c0e-8690-9a2edc0ffb8cCited by top-tier papers5
- CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding SpaceSohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun OhCVPR 2026 · 2 citations
- Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via LanguageNam Hyeon-Woo, Yebin Moon, Sohwi Lim, Kwon Byung-Ki et al.ICML 2026
- Necessary Conditions for Compositional Generalization of Embedding ModelsArnas Uselis, Andrea Dittadi, Seong Joon OhICML 2026
- Does Data Scaling Lead to Visual Compositional Generalization?Arnas Uselis, Andrea Dittadi, Seong Joon OhICML 2025
- Semantic Robustness Certification for Vision-Language ModelsPeiyu Yang, Paul MONTAGUE, Feng Liu, Andrew C. Cullen et al.ICML 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal RegressionWanhua Li, Xiaoke Huang, Zheng Zhu, Yansong Tang et al.NeurIPS 2022 · 65 citations
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image RetrievalSiting Li, Xiang Gao, Simon S. DuNeurIPS 2025 · 5 citations
- A View From Somewhere: Human-Centric Face RepresentationsJerone Theodore Alexander Andrews, Przemyslaw Joniak, Alice XiangICLR 2023 · 2 citations
- Self-Supervised Enhancement of Latent Discovery in GANsAdarsh Kappiyath, Silpa Vadakkeeveetil Sreelatha, S. SumitraAAAI 2022 · 3 citations
- Ranking-aware adapter for text-driven image ordering with CLIPWei-Hsiang Yu, Yen-Yu Lin, Ming-Hsuan Yang, Yi-Hsuan TsaiICLR 2025
