EIFFEL: a novel benchmark to measure bias of English heavy training on French idiomatic expressions
Charlotte Noel, Nicholas Asher, Olivier Gouvert, Farah Benamara, Julie Hunter
Abstract
Mainstream multilingual LLMs are generally trained on a much higher proportion of English than multilingual data, raising questions about their ability to capture linguistic features particular to non-English languages or to capture information important to non-anglophone cultures. We add to a growing effort to increase multilingual sensitivity in LLMs by developing a benchmark, EIFFEL, testing mastery of French idiomatic expressions in context. We fully explain the methodology, which exploits input from native French speakers, to make it reproducible for other languages. We compare mainstream multilingual LLMs with Frenchfocused LLMs both on standard LLM benchmarks and EIFFEL; EIFFEL brings out the benefits of higher proportions of French data and shows limitations of standard benchmarks for measuring multilingual competence. We also train from scratch a series of 1B SLMs with different proportions of French and English pretraining data that confirm EIFFEL's lessons. * En: "To burn bridges" * Distractor = "brûler" (lit. 'to burn") • To create the third distractor, select an English word similar to the English translation of the second distractor and translate the former into French. Examples: -"les pommes et les oranges" * En: "apples and oranges" * Distractor = "les pommes et les poires" (lit. "the apples and the pears") -"Brûler" * En: "To burn" * Distractor: "cramer" (lit. "to burn", "to torch") Step 4: Context generation Manually create a one-sentence, natural context for e, encourages an idiomatic, rather than literal, interpretation of e.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74b30200-cc8d-4946-aa87-7a3e639710b4Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Translate Meanings, Not Just Words: IdiomKB's Role in Optimizing Idiomatic Translation with Language ModelsShuang Li, Jiangjie Chen, Siyu Yuan, Xinyi Wu et al.AAAI 2024 · 44 citations
- The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language VariantsLucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe et al.ACL 2024 · 30 citations
- Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and EvaluationElaf Alhazmi, Quan Sheng, Wei Emma Zhang, Munazza Zaib et al.EMNLP 2024 · 15 citations
Related papers
- XToM: Exploring the Multilingual Theory of Mind for Large Language ModelsChunkit Chan, Yauwai Yim, Hongchuan Zeng, Zhiying Zou et al.ACL 2026
- Emergent Abilities of Large Language Models under Continued Pre-training for Language AdaptationAhmed Elhady, Eneko Agirre, Mikel ArtetxeACL 2025
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom AlignmentFengying Ye, Yanming Sun, Runzhe Zhan, Lidia S. Chao et al.ACL 2026 · 1 citation
- Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMsYanzhu Guo, Simone Conia, Zelin Zhou, Min Li et al.ACL 2025
- Understanding and Mitigating Language Confusion in LLMsKelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze et al.EMNLP 2024 · 10 citations
