Quantifying Prediction Consistency Under Fine-tuning Multiplicity in Tabular LLMs
Faisal Hamman, Pasan Dissanayake, Saumitra Mishra, Freddy Lécué, Sanghamitra Dutta
摘要
Fine-tuning LLMs on tabular classification tasks can lead to the phenomenon of fine-tuning multiplicity where equally well-performing models make conflicting predictions on the same input. Fine-tuning multiplicity can arise due to variations in the training process, e.g., seed, weight initialization, minor changes to training data, etc., raising concerns about the reliability of Tabular LLMs in high-stakes applications such as finance, hiring, education, healthcare. Our work formalizes this unique challenge of fine-tuning multiplicity in Tabular LLMs and proposes a novel measure to quantify the consistency of individual predictions without expensive model retraining. Our measure quantifies a prediction's consistency by analyzing (sampling) the model's local behavior around that input in the embedding space. Interestingly, we show that sampling in the local neighborhood can be leveraged to provide probabilistic guarantees on prediction consistency under a broad class of fine-tuned models, i.e., inputs with sufficiently high local stability (as defined by our measure) also remain consistent across several fine-tuned models with high probability. We perform experiments on multiple real-world datasets to show that our local stability measure preemptively captures consistency under actual multiplicity across several fine-tuned models, outperforming competing measures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction TuningYanjun Fu, Faisal Hamman, Sanghamitra DuttaNeurIPS 2025 · 被引用 15 次
- Few-Shot Knowledge Distillation of LLMs With Counterfactual ExplanationsFaisal Hamman, Pasan Dissanayake, Yanjun Fu, Sanghamitra DuttaNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper21
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
相关 Paper
- Consistent Counterfactuals for Deep ModelsEmily Black, Zifan Wang, Matt FredriksonICLR 2022 · 被引用 56 次
- Log Probability Tracking of LLM APIsTimothee Chauvin, Erwan Le Merrer, Francois Taiani, Gilles TredanICLR 2026 · 被引用 12 次
- Measuring the Instability of Fine-TuningYupei Du, Dong NguyenACL 2023 · 被引用 2 次
- Multiply Robust Estimation for Local Distribution Shifts with Multiple DomainsSteven Wilkins-Reeves, Xu Chen, Qi Ma, Christine Agarwal 等ICML 2024 · 被引用 2 次
- Implications of Model Indeterminacy for Explanations of Automated DecisionsMarc-Etienne Brunet, Ashton Anderson, Richard S. ZemelNeurIPS 2022 · 被引用 22 次
