ACL2026

EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressions

Charlotte Noel, Nicholas Asher, Olivier Gouvert, Farah Benamara, Julie Hunter

摘要

Mainstream multilingual LLMs are generally trained on a much higher proportion of English than multilingual data, raising questions about their ability to capture linguistic features particular to non-English languages or to capture information important to non-anglophone cultures. We add to a growing effort to increase multilingual sensitivity in LLMs by developing a benchmark, EIFFEL, testing mastery of French idiomatic expressions in context. We fully explain the methodology, which exploits input from native French speakers, to make it reproducible for other languages. We compare mainstream multilingual LLMs with Frenchfocused LLMs both on standard LLM benchmarks and EIFFEL; EIFFEL brings out the benefits of higher proportions of French data and shows limitations of standard benchmarks for measuring multilingual competence. We also train from scratch a series of 1B SLMs with different proportions of French and English pretraining data that confirm EIFFEL's lessons. * En: "To burn bridges" * Distractor = "brûler" (lit. 'to burn") • To create the third distractor, select an English word similar to the English translation of the second distractor and translate the former into French. Examples: -"les pommes et les oranges" * En: "apples and oranges" * Distractor = "les pommes et les poires" (lit. "the apples and the pears") -"Brûler" * En: "To burn" * Distractor: "cramer" (lit. "to burn", "to torch") Step 4: Context generation Manually create a one-sentence, natural context for e, encourages an idiomatic, rather than literal, interpretation of e.