Accelerating Benchmarking of Functional Connectivity Modeling via Structure-aware Core-set Selection
Ling Zhan, Zhen Li, Junjie Huang, Tao Jia
摘要
Benchmarking the hundreds of functional connectivity (FC) modeling methods on large-scale fMRI datasets is critical for reproducible neuroscience. However, the combinatorial explosion of model–data pairings makes exhaustive evaluation computationally prohibitive, preventing such assessments from becoming a routine pre-analysis step. To break this bottleneck, we reframe the challenge of FC benchmarking by selecting a small, representative core-set whose sole purpose is to preserve the relative performance ranking of FC operators. We formalize this as a ranking-preserving subset selection problem and propose Structure-aware Contrastive Learning for Core-set Selection (SCLCS), a self-supervised framework to select these core-sets. SCLCS first uses an adaptive Transformer to learn each sample's unique FC structure. It then introduces a novel Structural Perturbation Score (SPS) to quantify the stability of these learned structures during training, identifying samples that represent foundational connectivity archetypes. Finally, while SCLCS identifies stable samples via a top- ranking, we further introduce a density-balanced sampling strategy as a necessary correction to promote diversity, ensuring the final core-set is both structurally robust and distributionally representative. On the large-scale REST-meta-MDD dataset, SCLCS preserves the ground-truth model ranking with just 10% of the data, outperforming state-of-the-art (SOTA) core-set selection methods by up to 23.2% in ranking consistency (nDCG@k). To our knowledge, this is the first work to formalize core-set selection for FC operator benchmarking, thereby making large-scale operators comparisons a feasible and integral part of computational neuroscience. Code is publicly available on: https://github.com/lzhan94swu/SCLCS
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 被引用 398 次
相关 Paper
- Efficient Adversarial Contrastive Learning via Robustness-Aware Coreset SelectionXilie Xu, Jingfeng Zhang, Feng Liu, Masashi Sugiyama 等NeurIPS 2023 · 被引用 26 次
- Boosting Graph Contrastive Learning via Graph Contrastive SaliencyChunyu Wei, Yu Wang, Bing Bai, Kai Ni 等ICML 2023 · 被引用 31 次
- Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural EntropyJingyun Zhang, Hao Peng, Jianxin Li, Angsheng Li 等VLDB 2026
- An OpenMind for 3D Medical Vision Self-supervised LearningTassilo Wald, Constantin Ulrich, Jonathan Suprijadi, Sebastian Ziegler 等ICCV 2025 · 被引用 5 次
- Scaling Vision Transformers for Functional MRI with Flat MapsConnor Lane, Mihir Tripathy, Leema K Murali, Ratna Grandhi 等ICML 2026 · 被引用 3 次
