Lune

ICLR2023Top-tier venue

Plateau in Monotonic Linear Interpolation - A "Biased" View of Loss Landscape for Deep Networks

Xiang Wang, Annie N. Wang, Mo Zhou, Rong Ge

2023Year
3Top-tier citations

Abstract

Monotonic linear interpolation (MLI) -on the line connecting a random initialization with the minimizer it converges to, the loss and accuracy are monotonic -is a phenomenon that is commonly observed in the training of neural networks. Such a phenomenon may seem to suggest that optimization of neural networks is easy. In this paper, we show that the MLI property is not necessarily related to the hardness of optimization problems, and empirical observations on MLI for deep neural networks depend heavily on the biases. In particular, we show that interpolating both weights and biases linearly leads to very different influences on the final output, and when different classes have different last-layer biases on a deep network, there will be a long plateau in both the loss and accuracy interpolation (which existing theory of MLI cannot explain). We also show how the last-layer biases for different classes can be different even on a perfectly balanced dataset using a simple model. Empirically we demonstrate that similar intuitions hold on practical networks and realistic datasets.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b281a0b9-2189-46aa-aac9-8bb86d2e4b17

Cited by top-tier papers3

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines