Laplace Sample Information: Data Informativeness Through a Bayesian Lens
Johannes Kaiser, Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis
Abstract
Accurately estimating the informativeness of individual samples in a dataset is an important objective in deep learning, as it can guide sample selection, which can improve model efficiency and accuracy by removing redundant or potentially harmful samples. We propose Laplace Sample Information (LSI) measure of sample informativeness grounded in information theory widely applicable across model architectures and learning settings. LSI leverages a Bayesian approximation to the weight posterior and the KL divergence to measure the change in the parameter distribution induced by a sample of interest from the dataset. We experimentally show that LSI is effective in ordering the data with respect to typicality, detecting mislabeled samples, measuring class-wise informativeness, and assessing dataset difficulty. We demonstrate these capabilities of LSI on image and text data in supervised and unsupervised settings. Moreover, we show that LSI can be computed efficiently through probes and transfers well to the training of large models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00d767dc-5296-44f8-8470-1ca2786d52f3Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
Related papers
- Estimating informativeness of samples with Smooth Unique InformationHrayr Harutyunyan, Alessandro Achille, Giovanni Paolini, Orchid Majumder et al.ICLR 2021 · 26 citations
- Bayesian Deep Learning via Subnetwork InferenceErik A. Daxberger, Eric T. Nalisnick, James Urquhart Allingham, Javier Antorán et al.ICML 2021 · 108 citations
- Efficient Quantification of Multimodal Interaction at Sample LevelZequn Yang, Hongfa Wang, Di HuICML 2025
- Batch Loss Score for Dynamic Data PruningQing Zhou, Bingxuan Zhao, Tao Yang, Hongyuan Zhang et al.CVPR 2026
- Querying Easily Flip-flopped Samples for Deep Active LearningSeong Jin Cho, Gwangsu Kim, Junghyun Lee, Jinwoo Shin et al.ICLR 2024 · 8 citations
