Lune

KDD2026Top-tier venue

CaCuTe: Casual Cubic-Model Technique for Faster Optimization

Nazarii Tupitsa

2026Year

Abstract

We establish a local O(k−2)\mathcal{O}(k^{-2}) rate for the gradient update xk+1=xk−∇f(xk)/H∥∇f(xk)∥x^{k+1}=x^k-\nabla f(x^k)/\sqrt{H\|\nabla f(x^k)\|} under a 2H2H-Hessian--Lipschitz assumption. Regime detection relies on Hessian--vector products, avoiding Hessian formation or factorization. Incorporating this certificate into cubic-regularized Newton (CRN) and an accelerated variant enables per-iterate switching between the cubic and gradient steps while preserving CRN's global guarantees. The technique achieves the lowest wall-clock time among compared baselines in our experiments. In the first-order setting, the technique yields a monotone, adaptive, parameter-free method that inherits the local O(k−2)\mathcal{O}(k^{-2}) rate. Despite backtracking, the method shows superior wall-clock performance. Additionally, we cover smoothness relaxations beyond classical gradient--Lipschitzness, enabling tighter bounds, including global O(k−2)\mathcal{O}(k^{-2}) rates. Finally, we generalize the technique to the stochastic setting.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7cfdb993-e78b-4a00-8616-caea7b5dfa2c

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines