Deciphering the Extremes: A Novel Approach for Pathological Long-tailed Recognition in Scientific Discovery
Zhe Zhao, Haibin Wen, Xianfu Liu, Rui Mao, Pengkun Wang, Liheng Yu, Linjiang Chen, Bo An, Qingfu Zhang, Yang Wang
摘要
Scientific discovery across diverse fields increasingly grapples with datasets exhibiting pathological long-tailed distributions: a few common phenomena overshadow a multitude of rare yet scientifically critical instances. Unlike standard benchmarks, these scientific datasets often feature extreme imbalance coupled with a modest number of classes and limited overall sample volume, rendering existing long-tailed recognition (LTR) techniques ineffective. Such methods, biased by majority classes or prone to overfitting on scarce tail data, frequently fail to identify the very instances-novel materials, rare disease biomarkers, faint astronomical signals-that drive scientific breakthroughs. This paper introduces a novel, end-to-end framework explicitly designed to address pathological long-tailed recognition in scientific contexts. Our approach synergizes a Balanced Supervised Contrastive Learning (B-SCL) mechanism, which enhances the representation of tail classes by dynamically re-weighting their contributions, with a Smooth Objective Regularization (SOR) strategy that manages the inherent tension between tail-class focus and overall classification performance. We introduce and analyze the real-world ZincFluor chemical dataset (T = 137.54) and synthetic benchmarks with controllable extreme imbalances (CIFAR-LT variants). Extensive evaluations demonstrate our method's superior ability to decipher these extremes. Notably, on ZincFluor, our approach achieves a Tail Top-2 accuracy of 66.84%, significantly outperforming existing techniques. On CIFAR-10-LT with an imbalance ratio of 1000 (T = 100), our method achieves a tail-class accuracy of 38.99%, substantially leading the next best. These results underscore our framework's potential to unlock novel insights from complex, imbalanced scientific datasets, thereby accelerating discovery. We provide the detailed code in https://github.com/DataLab-atom/PLTR-SD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GoodDiffusion: Proactive Copyright Protection for Diffusion Generative Models via Learnable Sample-specific SignaturesShixi Qin, zhiyong yang, Shilong Bao, Zitai Wang 等ICML 2026
- Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of ViewJinping Wang, Zixin Tong, Zhiwu Xie, Zhiqiang GaoICML 2026
它引用的顶会 Paper9
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Parametric Contrastive LearningJiequan Cui, Zhisheng Zhong, Shu Liu, Bei Yu 等ICCV 2021 · 被引用 375 次
- Targeted Supervised Contrastive Learning for Long-Tailed RecognitionTianhong Li, Peng Cao, Yuan Yuan, Lijie Fan 等CVPR 2022 · 被引用 196 次
- Harnessing Hierarchical Label Distribution Variations in Test Agnostic Long-tail RecognitionZhiyong Yang, Qianqian Xu, Zitai Wang, Sicong Li 等ICML 2024 · 被引用 21 次
相关 Paper
- Balanced Contrastive Learning for Long-Tailed Visual RecognitionJianggang Zhu, Zheng Wang, Jingjing Chen, Yi-Ping Phoebe Chen 等CVPR 2022 · 被引用 194 次
- Subclass-balancing Contrastive Learning for Long-tailed RecognitionChengkai Hou, Jieyu Zhang, Haonan Wang, Tianyi ZhouICCV 2023 · 被引用 50 次
- Long-Tailed Recognition by Mutual Information Maximization between Latent Features and Ground-Truth LabelsMin-Kook Suh, Seung-Woo SeoICML 2023 · 被引用 30 次
- BCE3S: Binary Cross-Entropy Based Tripartite Synergistic Learning for Long-Tailed RecognitionWeijia Fan, Qiufu Li, Jiajun Wen, Xiaoyang PengAAAI 2026
- Exploring Balanced Feature Spaces for Representation LearningBingyi Kang, Yu Li, Sa Xie, Zehuan Yuan 等ICLR 2021 · 被引用 296 次
