Deciphering the Extremes: A Novel Approach for Pathological Long-tailed Recognition in Scientific Discovery
Zhe Zhao, Haibin Wen, Xianfu Liu, Rui Mao, Pengkun Wang, Liheng Yu, Linjiang Chen, Bo An, Qingfu Zhang, Yang Wang
Abstract
Scientific discovery across diverse fields increasingly grapples with datasets exhibiting pathological long-tailed distributions: a few common phenomena overshadow a multitude of rare yet scientifically critical instances. Unlike standard benchmarks, these scientific datasets often feature extreme imbalance coupled with a modest number of classes and limited overall sample volume, rendering existing long-tailed recognition (LTR) techniques ineffective. Such methods, biased by majority classes or prone to overfitting on scarce tail data, frequently fail to identify the very instances-novel materials, rare disease biomarkers, faint astronomical signals-that drive scientific breakthroughs. This paper introduces a novel, end-to-end framework explicitly designed to address pathological long-tailed recognition in scientific contexts. Our approach synergizes a Balanced Supervised Contrastive Learning (B-SCL) mechanism, which enhances the representation of tail classes by dynamically re-weighting their contributions, with a Smooth Objective Regularization (SOR) strategy that manages the inherent tension between tail-class focus and overall classification performance. We introduce and analyze the real-world ZincFluor chemical dataset (T = 137.54) and synthetic benchmarks with controllable extreme imbalances (CIFAR-LT variants). Extensive evaluations demonstrate our method's superior ability to decipher these extremes. Notably, on ZincFluor, our approach achieves a Tail Top-2 accuracy of 66.84%, significantly outperforming existing techniques. On CIFAR-10-LT with an imbalance ratio of 1000 (T = 100), our method achieves a tail-class accuracy of 38.99%, substantially leading the next best. These results underscore our framework's potential to unlock novel insights from complex, imbalanced scientific datasets, thereby accelerating discovery. We provide the detailed code in https://github.com/DataLab-atom/PLTR-SD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- GoodDiffusion: Proactive Copyright Protection for Diffusion Generative Models via Learnable Sample-specific SignaturesShixi Qin, zhiyong yang, Shilong Bao, Zitai Wang et al.ICML 2026
- Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of ViewJinping Wang, Zixin Tong, Zhiwu Xie, Zhiqiang GaoICML 2026
Builds on9
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Parametric Contrastive LearningJiequan Cui, Zhisheng Zhong, Shu Liu, Bei Yu et al.ICCV 2021 · 375 citations
- Targeted Supervised Contrastive Learning for Long-Tailed RecognitionTianhong Li, Peng Cao, Yuan Yuan, Lijie Fan et al.CVPR 2022 · 196 citations
- Harnessing Hierarchical Label Distribution Variations in Test Agnostic Long-tail RecognitionZhiyong Yang, Qianqian Xu, Zitai Wang, Sicong Li et al.ICML 2024 · 21 citations
Related papers
- Balanced Contrastive Learning for Long-Tailed Visual RecognitionJianggang Zhu, Zheng Wang, Jingjing Chen, Yi-Ping Phoebe Chen et al.CVPR 2022 · 194 citations
- Subclass-balancing Contrastive Learning for Long-tailed RecognitionChengkai Hou, Jieyu Zhang, Haonan Wang, Tianyi ZhouICCV 2023 · 50 citations
- Long-Tailed Recognition by Mutual Information Maximization between Latent Features and Ground-Truth LabelsMin-Kook Suh, Seung-Woo SeoICML 2023 · 30 citations
- BCE3S: Binary Cross-Entropy Based Tripartite Synergistic Learning for Long-Tailed RecognitionWeijia Fan, Qiufu Li, Jiajun Wen, Xiaoyang PengAAAI 2026
- Exploring Balanced Feature Spaces for Representation LearningBingyi Kang, Yu Li, Sa Xie, Zehuan Yuan et al.ICLR 2021 · 296 citations
