UniFast-HGR: Scalable and Efficient Maximal Correlation for Multimodal Models
Hongkang Zhang, Shao-Lun Huang, Yanlong Wang, Ercan KURUOGLU
摘要
This paper presents UniFast-HGR, a scalable surrogate for Hirschfeld-Gebelein-Rényi (HGR) maximal correlation in high-dimensional multimodal learning. The method replaces explicit covariance whitening with centered and ℓ 2normalized cosine alignment, uses the covarianceto-Gram trace identity to construct a local-batch structural surrogate, and removes invariant diagonal self-correlation through Trivial Spectrum Suppression (TSS). The resulting objective retains paired dependence maximization and the covariance-control role of Soft-HGR while replacing its finite-sample covariance estimator with a differentiable local-batch objective whose dominant structural cost is O(m 2 K) for local batch size m and feature dimension K. OptFast-HGR further reduces the practical memory burden by estimating the off-diagonal structural term through stochastic projection. Experiments across retrieval, image classification, remote sensing segmentation, and multimodal emotion recognition show consistent gains over covariance-based HGR/CCA variants and contrastive or neural MIestimator objectives on strong multimodal backbones, while microbenchmarks confirm stable behavior at extreme feature dimensions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 被引用 1,226 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
相关 Paper
- UniCon: Unified Framework for Efficient Contrastive Alignment via KernelsHangke Sui, Yuqing Wang, Minh N. DoICLR 2026 · 被引用 2 次
- SHGR: A Generalized Maximal Correlation CoefficientSamuel Stocksieker, Denys PommeretNeurIPS 2025
- Gramian Multimodal Representation Learning and AlignmentGiordano Cicchetti, Eleonora Grassucci, Luigi Sigillo, Danilo ComminielloICLR 2025
- Multi-Kernel Correlation-Attention Vision Transformer for Enhanced Contextual Understanding and Multi-Scale IntegrationHongkang Zhang, Shao-Lun Huang, Ercan E. Kuruoglu, Yanlong WangNeurIPS 2025
- Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph SpectrumWei Ai, Fuchen Zhang, Yuntao Shou, Tao Meng 等AAAI 2025 · 被引用 64 次
