Lune

NeurIPS2025顶会

On the rankability of visual embeddings

Ankit Sonthalia, Arnas Uselis, Seong Joon Oh

2025年份
4被引次数
5顶会引用

摘要

We study whether visual embedding models capture continuous, ordinal attributes along linear directions, which we term rank axes. We define a model as rankable for an attribute if projecting embeddings onto such an axis preserves the attribute's order. Across 7 popular encoders and 9 datasets with attributes like age, crowd count, head pose, aesthetics, and recency, we find that many embeddings are inherently rankable. Surprisingly, a small number of samples, or even just two extreme examples, often suffice to recover meaningful rank axes, without full-scale supervision. These findings open up new use cases for image ranking in vector databases and motivate further study into the structure and learning of rankable embeddings. Our code is available at https://github.com/aktsonthalia/rankable- vision-embeddings.

We examine two questions: (1) Are visual embeddings rankable? (2) How easily can we recover the rank axis for a given attribute? 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

To address (1), we evaluate 7 modern visual encoders, from ResNet to CLIP, across 9 datasets with 7 attributes: age, crowd count, 3 head pose angles (pitch, roll, yaw), image aesthetics, and recency. We find that many embedding spaces are indeed rankable (Section 3).

To address (2), we estimate the rank axis v A with minimal supervision. The structure of the embedding space makes full-dataset regression unnecessary. In many cases, a handful of annotated samples and, in some cases, a pair of samples x l (low) and x h (high) already recover non-trivial ranking performance. For the latter case, we define the rank axis as

This opens up the possibility for fast ordering of new images by arbitrary attributes. For example, a photo app lets users sort selfies by age appearance. It uses CLIP embeddings and two reference images: one of a child and one of an elderly person. The app computes v age without training. Users scroll from youngest-looking to oldest-looking faces in their album (Section 4).

Contributions:

  1. We define and motivate rankability as a property of visual embeddings, distinct from retrieval. 2. We study rankability across modern encoders and real-valued attributes; results show that current embeddings are rankable. 3. We show that rank axes can sometimes be recovered using only two or a handful of labelled samples.

2 Related work Embeddings for retrieval. Visual encoders are commonly used to index images in vector databases, enabling nearest neighbour search for retrieval tasks [42,52,33,57]. This setup, known as deep metric learning [5,6,42], predates vision-language models like CLIP [52]. CLIP and related models shifted the focus to cross-modal similarity modelling, where vision and language share a joint embedding space used for classification [52], retrieval [75,2], and captioning [39,26]. While the majority of work in visual encoders is devoted to the understanding of the local similarity structure, we study how visual embeddings support global ranking instead of just local retrieval.

Prior work has explored ways to improve the geometry of the embedding space. Order embeddings and hyperbolic representations have been used to model hierarchies [65,21,8,51]. Training disentangled representation [71] is considered critical for compositionality, where attributes are assigned to certain linear subspaces [60,3]. Others have defined concepts like uniformity and separability of the representations [70]. In this work, we focus on the analysis of a wide range of visual encoders, rather than introducing recipes for improvements.

A large body of work has examined the geometry and structure of CLIP's learned embedding space. CLIP and its derivatives have been studied extensively [4,53,32,77]. Several works have reported modality gaps between vision and language embeddings [12]. Some studies point to the absence of certain structures and capabilities in CLIP representations: attribute-object bindings [31,80,25], or the association of attributes to corresponding instances. Others argue that much information is already present in CLIP representations, including parts-of-speech and linguistic structure [44], attribute-object bindings [24], and compositional attributes [62,63]. The platonic representation hypothesis further suggests that models converge to similar internal structures [19]. In this work, we analyse the embedding geometry and structure for modern visual embeddings from the novel perspective of rankability.

Linearly probing an embedding. Linear probing is a fast and widely used method to test for the presence of concepts in visual embeddings [23,18,64]. It measures the accuracy of a linear classifier trained on intermediate-layer features, effectively testing whether a hyperplane can separate embeddings containing a concept from those that do not. This technique has been used to study the geometry of CLIP's embeddings [30] and to probe for specific information such as attribute-objec

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper5

问问它们各自怎么用它

它引用的顶会 Paper30

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖