Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language
Bo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yu Zhao, Yefeng Liu, Chenyu Zhu, Ruizhe Li, Jiahui Geng
Abstract
Instruction-following capability has become a major ability to be evaluated for Large Language Models (LLMs) (Brown et al., 2020; OpenAI, 2023; Bai et al., 2023) . However, existing datasets, such as IFEval (Zhou et al., 2023; Zeng et al., 2024) , are either predominantly monolingual and centered on English or simply machine translated to other languages, limiting their applicability in multilingual contexts. In this paper, we present an carefullycurated extension of IFEval to a localized multilingual version named Marco-Bench-MIF, covering 30 languages with varying levels of localization. Our benchmark addresses linguistic constraints (e.g., modifying capitalization requirements for Chinese) and cultural references (e.g., substituting region-specific company names in prompts) via a hybrid pipeline combining translation with verification. Through comprehensive evaluation of 20+ LLMs on our Marco-Bench-MIF, we found that: (1) 25-35% accuracy gap between high/low-resource languages, (2) model scales largely impact performance by 45-60% yet persists script-specific challenges, and (3) machine-translated data underestimates accuracy by 7-22% versus localized data. Our analysis identifies challenges in multilingual instruction following, including keyword consistency preservation and compositional constraint adherence across languages. Our Marco-Bench-MIF is available at https: //github.com/AIDC-AI/Marco-Bench-MIF .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37442a3a-a4a5-4a9a-987e-3f1e6fe3fe6eCited by top-tier papers1
Ask how each one uses itBuilds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Evaluating Large Language Models at Evaluating Instruction FollowingZhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng et al.ICLR 2024 · 299 citations
- Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAhmet Üstün, Viraat Aryabumi, Zheng Xin Yong, Wei-Yin Ko et al.ACL 2024
Related papers
- MaXIFE: Multilingual and Cross-lingual Instruction Following EvaluationYile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu et al.ACL 2025 · 5 citations
- FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong et al.ACL 2024 · 10 citations
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific TalksSara Papi, Maike Züfle, Marco Gaido, Beatrice Savoldi et al.ICLR 2026 · 20 citations
- MM-IFEngine: Towards Multimodal Instruction FollowingShengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang et al.ICCV 2025 · 3 citations
- McEval: Massively Multilingual Code EvaluationLinzheng Chai, Shukai Liu, Jian Yang, Yuwei Yin et al.ICLR 2025 · 1 citation
