Samples with Low Loss Curvature Improve Data Efficiency
Isha Garg, Kaushik Roy
Abstract
In this paper, we study the second order properties of the loss of trained deep neural networks with respect to the training data points to understand the curvature of the loss surface in the vicinity of these points. We find that there is an unexpected concentration of samples with very low curvature. We note that these low curvature samples are largely consistent across completely different architectures, and identifiable in the early epochs of training. We show that the curvature relates to the 'cleanliness' of the data points, with low curvatures samples corresponding to clean, higher clarity samples, representative of their category. Alternatively, high curvature samples are often occluded, have conflicting features and visually atypical of their category. Armed with this insight, we introduce SLo-Curves, a novel coreset identification and training algorithm. SLocurves identifies the samples with low curvatures as being more data-efficient and trains on them with an additional regularizer that penalizes high curvature of the loss surface in their vicinity. We demonstrate the efficacy of SLo-Curves on CIFAR-10 and CIFAR-100 datasets, where it outperforms state of the art coreset selection methods at small coreset sizes by up to 9%. The identified coresets generalize across architectures, and hence can be pre-computed to generate condensed versions of datasets for use in downstream tasks. Code is available at https://github.com/isha- garg/SLo-Curves.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3eba75dc-25d7-4780-885b-978735c27791Cited by top-tier papers7
- Memorization Through the Lens of Curvature of Loss Function Around SamplesIsha Garg, Deepak Ravikumar, Kaushik RoyICML 2024 · 25 citations
- Unveiling Privacy, Memorization, and Input Curvature LinksDeepak Ravikumar, Efstathia Soufleri, Abolfazl Hashemi, Kaushik RoyICML 2024 · 16 citations
- Curvature Clues: Decoding Deep Learning Privacy with Input Loss CurvatureDeepak Ravikumar, Efstathia Soufleri, Kaushik RoyNeurIPS 2024 · 10 citations
- Sharpness-diversity tradeoff: improving flat ensembles with SharpBalanceHaiquan Lu, Xiaotian Liu, Yefan Zhou, Qunli Li et al.NeurIPS 2024 · 4 citations
- Memorization Through the Lens of Sample GradientsDeepak Ravikumar, Efstathia Soufleri, Abolfazl Hashemi, Kaushik RoyICLR 2026
Builds on16
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
Related papers
- Stable Coresets via Posterior Sampling: Aligning Induced and Full Loss LandscapesWei-Kai Chang, Rajiv KhannaNeurIPS 2025
- Adaptive Second Order Coresets for Data-efficient Machine LearningOmead Pooladzandi, David Davini, Baharan MirzasoleimanICML 2022 · 83 citations
- Efficient Representativeness-Aware Coreset SelectionZihao Cheng, Binrui Wu, Zhiwei Li, Yuesen Liao et al.NeurIPS 2025 · 1 citation
- Towards Sustainable Learning: Coresets for Data-efficient Deep LearningYu Yang, Hao Kang, Baharan MirzasoleimanICML 2023 · 58 citations
- Efficient Core-set Selection for Deep Learning Through Squared Loss MinimizationJianting ChenICML 2025
