Lune

EMNLP2024顶会

Rethinking the Evaluation of In-Context Learning for LLMs

Guoxin Yu, Lemao Liu, Mo Yu, Yue Yu, Xiang Ao

2024年份

摘要

In-context learning (ICL) has demonstrated excellent performance across various downstream NLP tasks, especially when synergized with powerful large language models (LLMs). Existing studies evaluate ICL methods primarily based on downstream task performance. This evaluation protocol overlooks the significant cost associated with the demonstration configuration process, i.e., tuning the demonstration as the ICL prompt. However, in this work, we point out that the evaluation protocol leads to unfair comparisons and potentially biased evaluation, because we surprisingly find the correlation between the configuration costs and task performance. Then we call for a twodimensional evaluation paradigm that considers both of these aspects, facilitating a fairer comparison. Finally, based on our empirical finding that the optimized demonstration on one language model generalizes across language models of different sizes, we introduce a simple yet efficient strategy that can be applied to any ICL method as a plugin, yielding a better trade-off between the two dimensions according to the proposed evaluation paradigm.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 9e4b9665-ce6b-428a-8ed2-0a12145e87c8

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖