Multi-modal Extreme Classification
Anshul Mittal, Kunal Dahiya, Shreya Malani, Janani Ramaswamy, Seba Ann Kuruvilla, Jitendra Ajmera, Keng-hao Chang, Sumeet Agarwal, Purushottam Kar, Manik Varma
摘要
This paper develops the MUFIN technique for extreme classification (XC) tasks with millions of labels where data-points and labels are endowed with visual and textual de-scriptors. Applications of MUFIN to product-to-product recommendation and bid query prediction over several mil-lions of products are presented. Contemporary multi-modal methods frequently rely on purely embedding-based meth-ods. On the other hand, XC methods utilize classifier ar-chitectures to offer superior accuracies than embedding-only methods but mostly focus on text-based categorization tasks. MUFIN bridges this gap by reformulating multi-modal categorization as an XC problem with several mil-lions of labels. This presents the twin challenges of devel-oping multi-modal architectures that can offer embeddings sufficiently expressive to allow accurate categorization over millions of labels; and training and inference routines that scale logarithmically in the number of labels. MUFIN de-velops an architecture based on cross-modal attention and trains it in a modular fashion using pre-training and positive and negative mining. A novel product-to-product rec-ommendation dataset MM-AmazonTitles-300K containing over 300K products was curated from publicly available amazon.com listings with each product endowed with a title and multiple images. On the MM-AmazonTitles-300K and Polyvore datasets, and a dataset with over 4 million labels curated from click logs of the Bing search engine, MUFIN offered at least 3% higher accuracy than leading text-based, image-based and multi-modal techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Dual-Encoders for Extreme Multi-label ClassificationNilesh Gupta, Devvrit, Ankit Singh Rawat, Srinadh Bhojanapalli 等ICLR 2024 · 被引用 8 次
- Deep Encoders with Auxiliary Parameters for Extreme ClassificationKunal Dahiya, Sachin Yadav, Sushant Sondhi, Deepak Saini 等KDD 2023 · 被引用 6 次
- Deep Fuzzy Multi-view Learning for Reliable ClassificationSiyuan Duan, Yuan Sun, Dezhong Peng, Guiduo Duan 等ICML 2025
- Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal FrameworkDiego Ortego, Marlon Rodríguez, Mario Almagro, Kunal Dahiya 等AAAI 2026
- A Generative Approach for Wikipedia-Scale Visual Entity RecognitionMathilde Caron, Ahmet Iscen, Alireza Fathi, Cordelia SchmidCVPR 2024
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang 等AAAI 2021 · 被引用 170 次
相关 Paper
- GalaXC: Graph Neural Networks with Labelwise Attention for Extreme ClassificationDeepak Saini, Arnav Kumar Jain, Kushal Dave, Jian Jiao 等WWW 2021 · 被引用 49 次
- SiameseXML: Siamese Networks meet Extreme Classifiers with 100M LabelsKunal Dahiya, Ananye Agarwal, Deepak Saini, Gururaj K 等ICML 2021 · 被引用 61 次
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 被引用 147 次
- ECLARE: Extreme Classification with Label Graph CorrelationsAnshul Mittal, Noveen Sachdeva, Sheshansh Agrawal, Sumeet Agarwal 等WWW 2021 · 被引用 71 次
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 被引用 4 次
