Multi-modal Extreme Classification
Anshul Mittal, Kunal Dahiya, Shreya Malani, Janani Ramaswamy, Seba Ann Kuruvilla, Jitendra Ajmera, Keng-hao Chang, Sumeet Agarwal, Purushottam Kar, Manik Varma
Abstract
This paper develops the MUFIN technique for extreme classification (XC) tasks with millions of labels where data-points and labels are endowed with visual and textual de-scriptors. Applications of MUFIN to product-to-product recommendation and bid query prediction over several mil-lions of products are presented. Contemporary multi-modal methods frequently rely on purely embedding-based meth-ods. On the other hand, XC methods utilize classifier ar-chitectures to offer superior accuracies than embedding-only methods but mostly focus on text-based categorization tasks. MUFIN bridges this gap by reformulating multi-modal categorization as an XC problem with several mil-lions of labels. This presents the twin challenges of devel-oping multi-modal architectures that can offer embeddings sufficiently expressive to allow accurate categorization over millions of labels; and training and inference routines that scale logarithmically in the number of labels. MUFIN de-velops an architecture based on cross-modal attention and trains it in a modular fashion using pre-training and positive and negative mining. A novel product-to-product rec-ommendation dataset MM-AmazonTitles-300K containing over 300K products was curated from publicly available amazon.com listings with each product endowed with a title and multiple images. On the MM-AmazonTitles-300K and Polyvore datasets, and a dataset with over 4 million labels curated from click logs of the Bing search engine, MUFIN offered at least 3% higher accuracy than leading text-based, image-based and multi-modal techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c8be441-9356-4e91-8f4d-36df73ddc2bfCited by top-tier papers5
- Dual-Encoders for Extreme Multi-label ClassificationNilesh Gupta, Devvrit, Ankit Singh Rawat, Srinadh Bhojanapalli et al.ICLR 2024 · 8 citations
- Deep Encoders with Auxiliary Parameters for Extreme ClassificationKunal Dahiya, Sachin Yadav, Sushant Sondhi, Deepak Saini et al.KDD 2023 · 6 citations
- Deep Fuzzy Multi-view Learning for Reliable ClassificationSiyuan Duan, Yuan Sun, Dezhong Peng, Guiduo Duan et al.ICML 2025
- Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal FrameworkDiego Ortego, Marlon Rodríguez, Mario Almagro, Kunal Dahiya et al.AAAI 2026
- A Generative Approach for Wikipedia-Scale Visual Entity RecognitionMathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia SchmidCVPR 2024
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang et al.AAAI 2021 · 170 citations
Related papers
- GalaXC: Graph Neural Networks with Labelwise Attention for Extreme ClassificationDeepak Saini, Arnav Kumar Jain, Kushal Dave, Jian Jiao et al.WWW 2021 · 49 citations
- SiameseXML: Siamese Networks meet Extreme Classifiers with 100M LabelsKunal Dahiya, Ananye Agarwal, Deepak Saini, Gururaj K et al.ICML 2021 · 61 citations
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 147 citations
- ECLARE: Extreme Classification with Label Graph CorrelationsAnshul Mittal, Noveen Sachdeva, Sheshansh Agrawal, Sumeet Agarwal et al.WWW 2021 · 71 citations
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 4 citations
