Lune

EMNLP2024Top-tier venue

Rethinking the Evaluation of In-Context Learning for LLMs

Guoxin Yu, Lemao Liu, Mo Yu, Yue Yu, Xiang Ao

2024Year

Abstract

In-context learning (ICL) has demonstrated excellent performance across various downstream NLP tasks, especially when synergized with powerful large language models (LLMs). Existing studies evaluate ICL methods primarily based on downstream task performance. This evaluation protocol overlooks the significant cost associated with the demonstration configuration process, i.e., tuning the demonstration as the ICL prompt. However, in this work, we point out that the evaluation protocol leads to unfair comparisons and potentially biased evaluation, because we surprisingly find the correlation between the configuration costs and task performance. Then we call for a twodimensional evaluation paradigm that considers both of these aspects, facilitating a fairer comparison. Finally, based on our empirical finding that the optimized demonstration on one language model generalizes across language models of different sizes, we introduce a simple yet efficient strategy that can be applied to any ICL method as a plugin, yielding a better trade-off between the two dimensions according to the proposed evaluation paradigm.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9e4b9665-ce6b-428a-8ed2-0a12145e87c8

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines