Lune

NeurIPS2025Top-tier venue

On the rankability of visual embeddings

Ankit Sonthalia, Arnas Uselis, Seong Joon Oh

2025Year
4Citations
5Top-tier citations

Abstract

We study whether visual embedding models capture continuous, ordinal attributes along linear directions, which we term rank axes. We define a model as rankable for an attribute if projecting embeddings onto such an axis preserves the attribute's order. Across 7 popular encoders and 9 datasets with attributes like age, crowd count, head pose, aesthetics, and recency, we find that many embeddings are inherently rankable. Surprisingly, a small number of samples, or even just two extreme examples, often suffice to recover meaningful rank axes, without full-scale supervision. These findings open up new use cases for image ranking in vector databases and motivate further study into the structure and learning of rankable embeddings. Our code is available at https://github.com/aktsonthalia/rankable- vision-embeddings.

We examine two questions: (1) Are visual embeddings rankable? (2) How easily can we recover the rank axis for a given attribute? 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

To address (1), we evaluate 7 modern visual encoders, from ResNet to CLIP, across 9 datasets with 7 attributes: age, crowd count, 3 head pose angles (pitch, roll, yaw), image aesthetics, and recency. We find that many embedding spaces are indeed rankable (Section 3).

To address (2), we estimate the rank axis v A with minimal supervision. The structure of the embedding space makes full-dataset regression unnecessary. In many cases, a handful of annotated samples and, in some cases, a pair of samples x l (low) and x h (high) already recover non-trivial ranking performance. For the latter case, we define the rank axis as

This opens up the possibility for fast ordering of new images by arbitrary attributes. For example, a photo app lets users sort selfies by age appearance. It uses CLIP embeddings and two reference images: one of a child and one of an elderly person. The app computes v age without training. Users scroll from youngest-looking to oldest-looking faces in their album (Section 4).

Contributions:

  1. We define and motivate rankability as a property of visual embeddings, distinct from retrieval. 2. We study rankability across modern encoders and real-valued attributes; results show that current embeddings are rankable. 3. We show that rank axes can sometimes be recovered using only two or a handful of labelled samples.

2 Related work Embeddings for retrieval. Visual encoders are commonly used to index images in vector databases, enabling nearest neighbour search for retrieval tasks [42,52,33,57]. This setup, known as deep metric learning [5,6,42], predates vision-language models like CLIP [52]. CLIP and related models shifted the focus to cross-modal similarity modelling, where vision and language share a joint embedding space used for classification [52], retrieval [75,2], and captioning [39,26]. While the majority of work in visual encoders is devoted to the understanding of the local similarity structure, we study how visual embeddings support global ranking instead of just local retrieval.

Prior work has explored ways to improve the geometry of the embedding space. Order embeddings and hyperbolic representations have been used to model hierarchies [65,21,8,51]. Training disentangled representation [71] is considered critical for compositionality, where attributes are assigned to certain linear subspaces [60,3]. Others have defined concepts like uniformity and separability of the representations [70]. In this work, we focus on the analysis of a wide range of visual encoders, rather than introducing recipes for improvements.

A large body of work has examined the geometry and structure of CLIP's learned embedding space. CLIP and its derivatives have been studied extensively [4,53,32,77]. Several works have reported modality gaps between vision and language embeddings [12]. Some studies point to the absence of certain structures and capabilities in CLIP representations: attribute-object bindings [31,80,25], or the association of attributes to corresponding instances. Others argue that much information is already present in CLIP representations, including parts-of-speech and linguistic structure [44], attribute-object bindings [24], and compositional attributes [62,63]. The platonic representation hypothesis further suggests that models converge to similar internal structures [19]. In this work, we analyse the embedding geometry and structure for modern visual embeddings from the novel perspective of rankability.

Linearly probing an embedding. Linear probing is a fast and widely used method to test for the presence of concepts in visual embeddings [23,18,64]. It measures the accuracy of a linear classifier trained on intermediate-layer features, effectively testing whether a hyperplane can separate embeddings containing a concept from those that do not. This technique has been used to study the geometry of CLIP's embeddings [30] and to probe for specific information such as attribute-objec

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c89bfa96-982e-4c0e-8690-9a2edc0ffb8c

Cited by top-tier papers5

Ask how each one uses it

Builds on30

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines