Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
Zheng Luo, Thirulogasankar Pranav Kutralingam, Ogochukwu N. Okoani, Wanpeng Xu, Hua Wei, Xiyang Hu
Abstract
Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent work reports strong tool-calling performance under standard English-centric evaluations, the robustness of tool calling under multilingual user interactions remains underexplored. In this work, we introduce MLCL, a diagnostic benchmark, and conduct a systematic evaluation of multilingual tool calling across Chinese, Hindi, and the low-resource language Igbo. Through fine-grained error analysis, we show that many failures occur despite correct intent understanding and tool selection. We identify parameter value language mismatch as a dominant failure mode, where models generate semantically appropriate parameter values in the user's language, violating language-invariant execution conventions. We further evaluate several inference-time system strategies and find that while these strategies substantially reduce language-induced execution errors, none of them can fully recover English-level performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a8e9a97-370b-4c59-8610-d9cbdb00beddCited by top-tier papers1
Ask how each one uses itBuilds on10
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang et al.ICLR 2026 · 406 citations
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng et al.EMNLP 2023 · 91 citations
Related papers
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare QueriesYiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu et al.WWW 2024 · 126 citations
- OrchestrationBench: LLM-Driven Agentic Planning and Tool Use in Multi-Domain ScenariosAelim Ahn, Sooyeon Lee, Hyosun Wang, Chiwan Park et al.ICLR 2026
- INCLUDE: Evaluating Multilingual Language Understanding with Regional KnowledgeAngelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen et al.ICLR 2025
- CrossPL: Systematic Evaluation of Large Language Models for Cross Programming Language Interoperating Code Generationzhanhang xiong, Dongxia Wang, Yuekang Li, Xinyuan An et al.ICLR 2026
- Benchmarking LLM Tool-Use in the WildPeijie Yu, Wei Liu, Yifan Yang, Jinjian Li et al.ICLR 2026 · 20 citations
