The Performance Analysis of Generalized Margin Maximizers on Separable Data
Fariborz Salehi, Ehsan Abbasi, Babak Hassibi
Abstract
Logistic models are commonly used for binary classification tasks. The success of such models has often been attributed to their connection to maximum-likelihood estimators. It has been shown that gradient descent algorithm, when applied on the logistic loss, converges to the max-margin classifier (a.k.a. hard-margin SVM). The performance of the max-margin classifier has been recently analyzed in (Montanari et al., 2019; Deng et al., 2019) . Inspired by these results, in this paper, we present and study a more general setting, where the underlying parameters of the logistic model possess certain structures (sparse, block-sparse, low-rank, etc.) and introduce a more general framework (which is referred to as "Generalized Margin Maximizer", GMM). While classical max-margin classifiers minimize the 2-norm of the parameter vector subject to linearly separating the data, GMM minimizes any arbitrary convex function of the parameter vector. We provide a precise analysis of the performance of GMM via the solution of a system of nonlinear equations. We also provide a detailed study for three special cases: (1) ℓ 2 -GMM that is the max-margin classifier, (2) ℓ 1 -GMM which encourages sparsity, and (3) ℓ ∞ -GMM which is often used when the parameter vector has binary entries. Our theoretical results are validated by extensive simulation results across a range of parameter values, problem instances, and model structures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c996b8a-ff82-420d-a127-1e69e5a35850Cited by top-tier papers6
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural NetworksXiangyu Chang, Yingcong Li, Samet Oymak, Christos ThrampoulidisAAAI 2021 · 58 citations
- Benign Overfitting in Multiclass Classification: All Roads Lead to InterpolationKe Wang, Vidya Muthukumar, Christos ThrampoulidisNeurIPS 2021 · 56 citations
- Error Analysis of Spherically Constrained Least Squares Reformulation in Solving the Stackelberg Prediction GameXiyuan Li, Weiwei LiuNeurIPS 2024 · 1 citation
- The Reliability of OKRidge Method in Solving Sparse Ridge Regression ProblemsXiyuan Li, Youjun Wang, Weiwei LiuNeurIPS 2024
Related papers
- Learning Gaussian Mixtures with Generalized Linear Models: Precise Asymptotics in High-dimensionsBruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco et al.NeurIPS 2021 · 70 citations
- Does Momentum Change the Implicit Regularization on Separable Data?Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun et al.NeurIPS 2022 · 29 citations
- The Implicit Bias of Adam on Separable DataChenyang Zhang, Difan Zou, Yuan CaoNeurIPS 2024 · 37 citations
- A Discriminative Gaussian Mixture Model with SparsityHideaki Hayashi, Seiichi UchidaICLR 2021 · 8 citations
- Mirror Descent Maximizes Generalized Margin and Can Be Implemented EfficientlyHaoyuan Sun, Kwangjun Ahn, Christos Thrampoulidis, Navid AzizanNeurIPS 2022 · 33 citations
