Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification
Rui Wang, Peipei Li, Huaibo Huang, Chunshui Cao, Ran He, Zhaofeng He
摘要
We present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspire us to leverage the rich ordinal priors in human language by converting the original task into a visionlanguage alignment task. Consequently, we propose L2RCLIP, which fully utilizes the language priors from two perspectives. First, we introduce a complementary prompt tuning technique called RankFormer, designed to enhance the ordering relation of original rank prompts. It employs token-level attention with residual-style prompt blending in the word embedding space. Second, to further incorporate language priors, we revisit the approximate bound optimization of vanilla cross-entropy loss and restructure it within the cross-modal embedding space. Consequently, we propose a cross-modal ordinal pairwise loss to refine the CLIP feature space, where texts and images maintain both semantic alignment and ordering alignment. Extensive experiments on three ordinal classification tasks, including facial age estimation, historical color image (HCI) classification, and aesthetic assessment demonstrate its promising performance. The code is available at https://github.com/raywang335/L2RCLIP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- On the rankability of visual embeddingsAnkit Sonthalia, Arnas Uselis, Seong Joon OhNeurIPS 2025 · 被引用 4 次
- ChatEdit: Towards Multi-turn Interactive Facial Image Editing via DialogueXing Cui, Zekun Li, Pei Li, Yibo Hu 等EMNLP 2023 · 被引用 4 次
- TSPRank: Bridging Pairwise and Listwise Methods with a Bilinear Travelling Salesman ModelWeixian Waylon Li, Yftah Ziser, Yifei Xie, Shay B. Cohen 等KDD 2025 · 被引用 2 次
- DiffoR: A Unified Continuous Generative Framework for Universal Ordinal RegressionHongxu Ma, Lin Wang, Chenghou Jin, Han Zhou 等KDD 2026 · 被引用 1 次
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 被引用 1 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
相关 Paper
- OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal RegressionWanhua Li, Xiaoke Huang, Zheng Zhu, Yansong Tang 等NeurIPS 2022 · 被引用 65 次
- Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via LanguageNam Hyeon-Woo, Yebin Moon, Sohwi Lim, Kwon Byung-Ki 等ICML 2026
- Ranking-aware adapter for text-driven image ordering with CLIPWei-Hsiang Yu, Yen-Yu Lin, Ming-Hsuan Yang, Yi-Hsuan TsaiICLR 2025
- LAMM: Label Alignment for Multi-Modal Prompt LearningJingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu 等AAAI 2024 · 被引用 33 次
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng 等ICCV 2025 · 被引用 1 次
