Dual-Encoders for Extreme Multi-label Classification
Nilesh Gupta, Devvrit, Ankit Singh Rawat, Srinadh Bhojanapalli, Prateek Jain, Inderjit S. Dhillon
摘要
Dual-encoder (DE) models are widely used in retrieval tasks, most commonly studied on open QA benchmarks that are often characterized by multi-class and limited training data. In contrast, their performance in multi-label and data-rich retrieval settings like extreme multi-label classification (XMC), remains under-explored. Current empirical evidence indicates that DE models fall significantly short on XMC benchmarks, where SOTA methods linearly scale the number of learnable parameters with the total number of classes (documents in the corpus) by employing per-class classification head. To this end, we first study and highlight that existing multi-label contrastive training losses are not appropriate for training DE models on XMC tasks. We propose decoupled softmax loss - a simple modification to the InfoNCE loss - that overcomes the limitations of existing contrastive losses. We further extend our loss design to a soft top-k operator-based loss which is tailored to optimize top-k prediction performance. When trained with our proposed loss functions, standard DE models alone can match or outperform SOTA methods by up to 2% at Precision@1 even on the largest XMC datasets while being 20x smaller in terms of the number of trainable parameters. This leads to more parameter-efficient and universally applicable solutions for retrieval tasks. Our code and models are publicly available at https://github.com/nilesh2797/dexml.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Gandalf: Learning Label-label Correlations in Extreme Multi-label Classification via Label FeaturesSiddhant Kharbanda, Devaansh Gupta, Erik Schultheis, Atmadeep Banerjee 等KDD 2024 · 被引用 6 次
- Sampled Estimators For Softmax Must Be BiasedLi-Chung Lin, Yaxu Liu, Chih-Jen LinNeurIPS 2025 · 被引用 3 次
- UniDEC : Unified Dual Encoder and Classifier Training for Extreme Multi-Label ClassificationSiddhant Kharbanda, Devaansh Gupta, Gururaj K, Pankaj Malhotra 等WWW 2025 · 被引用 2 次
- On the Necessity of World Knowledge for Mitigating Missing Labels in Extreme ClassificationJatin Prakash, Anirudh Buvanesh, Bishal Santra, Deepak Saini 等KDD 2025 · 被引用 1 次
- Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal FrameworkDiego Ortego, Marlon Rodríguez, Mario Almagro, Kunal Dahiya 等AAAI 2026
它引用的顶会 Paper17
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang 等AAAI 2021 · 被引用 170 次
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 被引用 147 次
相关 Paper
- Deep Encoders with Auxiliary Parameters for Extreme ClassificationKunal Dahiya, Sachin Yadav, Sushant Sondhi, Deepak Saini 等KDD 2023 · 被引用 6 次
- A Gradient Accumulation Method for Dense Retriever under Memory ConstraintJaehee Kim, Yukyung Lee, Pilsung KangNeurIPS 2024 · 被引用 10 次
- Convex Surrogates for Unbiased Loss Functions in Extreme Classification With Missing LabelsMohammadreza Qaraei, Erik Schultheis, Priyanshu Gupta, Rohit BabbarWWW 2021 · 被引用 29 次
- Differentiable Top-k Classification LearningFelix Petersen, Hilde Kuehne, Christian Borgelt, Oliver DeussenICML 2022 · 被引用 48 次
- Two-Way Multi-Label LossTakumi KobayashiCVPR 2023
