DA-Pred: Performance Prediction for Text Summarization under Domain-Shift and Instruct-Tuning
Anum Afzal, Florian Matthes, Alexander R. Fabbri
摘要
Large Language Models (LLMs) often don’t perform as expected under Domain Shift or after Instruct-tuning. A reliable indicator of LLM performance in these settings could assist in decision-making. We present a method that uses the known performance in high-resource domains and fine-tuning settings to predict performance in low-resource domains or base models, respectively. In our paper, we formulate the task of performance prediction, construct a dataset for it, and train regression models to predict the said change in performance. Our proposed methodology is lightweight and, in practice, can help researchers & practitioners decide if resources should be allocated for data labeling and LLM Instruct-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- Predicting Performance for Natural Language Processing TasksMengzhou Xia, Antonios Anastasopoulos, Ruochen Xu, Yiming Yang 等ACL 2020 · 被引用 39 次
- Making Science Simple: Corpora for the Lay Summarisation of Scientific LiteratureTomas Goldsack, Zhihao Zhang, Chenghua Lin, Carolina ScartonEMNLP 2022 · 被引用 38 次
相关 Paper
- TuneAhead: Predicting Fine-tuning Performance Before Training BeginsYuxiang Luo, Haonan Long, Chen Wang, Qiqi Duan 等ICML 2026
- WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in ThaiPeerat Limkonchotiwat, Pume Tuchinda, Lalita Lowphansirikul, Surapon Nonesung 等EMNLP 2025
- SAE as a Crystal Ball: Interpretable Features Predict Cross-domain Transferability of LLMs without TrainingQi Zhang, Yifei Wang, Xiaohan Wang, Jiajun Chai 等ICLR 2026 · 被引用 1 次
- Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsJiuding Sun, Chantal Shaib, Byron C. WallaceICLR 2024 · 被引用 75 次
- Predicting Fine-Tuning Performance with ProbingZining Zhu, Soroosh Shahtalebi, Frank RudziczEMNLP 2022 · 被引用 6 次
