SWIFT: Scalable Wasserstein Factorization for Sparse Nonnegative Tensors
Ardavan Afshar, Kejing Yin, Sherry Yan, Cheng Qian, Joyce C. Ho, Haesun Park, Jimeng Sun
摘要
Existing tensor factorization methods assume that the input tensor follows some specific distribution (i.e. Poisson, Bernoulli, and Gaussian), and solve the factorization by minimizing some empirical loss functions defined based on the corresponding distribution. However, it suffers from several drawbacks: 1) In reality, the underlying distributions are complicated and unknown, making it infeasible to be approximated by a simple distribution. 2) The correlation across dimensions of the input tensor is not well utilized, leading to sub-optimal performance. Although heuristics were proposed to incorporate such correlation as side information under Gaussian distribution, they can not easily be generalized to other distributions. Thus, a more principled way of utilizing the correlation in tensor factorization models is still an open challenge. Without assuming any explicit distribution, we formulate the tensor factorization as an optimal transport problem with Wasserstein distance, which can handle non-negative inputs.
We introduce SWIFT, which minimizes the Wasserstein distance that measures the distance between the input tensor and that of the reconstruction. In particular, we define the N-th order tensor Wasserstein loss for the widely used tensor CP factorization and derive the optimization algorithm that minimizes it. By leveraging sparsity structure and different equivalent formulations for optimizing computational efficiency, SWIFT is as scalable as other well-known CP algorithms. Using the factor matrices as features, SWIFT achieves up to 9.65% and 11.31% relative improvement over baselines for downstream prediction tasks. Under the noisy conditions, SWIFT achieves up to 15% and 17% relative improvements over the best competitors for the prediction tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Gromov-Wasserstein Factorization Models for Graph ClusteringHongteng XuAAAI 2020 · 被引用 56 次
- LogPar: Logistic PARAFAC2 Factorization for Temporal Binary Data with Missing ValuesKejing Yin, Ardavan Afshar, Joyce C. Ho, William K. Cheung 等KDD 2020 · 被引用 33 次
- Beyond Rank-1: Discovering Rich Community Structure in Multi-Aspect GraphsEkta Gujral, Ravdeep Pasricha, Evangelos E. PapalexakisWWW 2020 · 被引用 24 次
相关 Paper
- Optimal Tensor TransportTanguy Kerdoncuff, Rémi Emonet, Michaël Perrot, Marc SebbanAAAI 2022 · 被引用 3 次
- Provable Online CP/PARAFAC Decomposition of a Structured Tensor via Dictionary LearningSirisha Rambhatla, Xingguo Li, Jarvis D. HauptNeurIPS 2020 · 被引用 13 次
- Uncertainty quantification for nonconvex tensor completion: Confidence intervals, heteroscedasticity and optimalityChangxiao Cai, H. Vincent Poor, Yuxin ChenICML 2020 · 被引用 26 次
- Score-Based Model for Low-Rank Tensor RecoveryZhengyun Cheng, Changhao Wang, Guanwen Zhang, Yi Xu 等AAAI 2026
- Transforms based Tensor Robust PCA: Corrupted Low-Rank Tensors Recovery via Convex OptimizationCanyi LuICCV 2021 · 被引用 29 次
