FilBench: Can LLMs Understand and Generate Filipino?
Lester James Validad Miranda, Elyanah Aco, Conner G. Manuel, Jan Christian Blaise Cruz, Joseph Marvin Imperial
Abstract
Despite the impressive performance of LLMs on English-based tasks, little is known about their capabilities in specific languages such as Filipino. In this work, we address this gap by introducing FILBENCH, a Filipino-centric benchmark designed to evaluate LLMs across a diverse set of tasks and capabilities in Filipino, Tagalog, and Cebuano. We carefully curate the tasks in FILBENCH to reflect the priorities and trends of NLP research in the Philippines such as Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. By evaluating 27 state-of-the-art LLMs on FILBENCH, we find that several LLMs suffer from reading comprehension and translation capabilities. Our results indicate that FILBENCH is challenging, with the best model, GPT-4o, achieving only a score of 72.23%. Moreover, we also find that models trained specifically for Southeast Asian languages tend to underperform on FIL-BENCH, with the highest-performing model, SEA-LION v3 70B, achieving only a score of 61.07%. Our work demonstrates the value of curating language-specific LLM benchmarks to aid in driving progress on Filipino NLP and increasing the inclusion of Philippine languages in LLM development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72ee402d-3cbd-48ed-ad1b-8b4c89347b81Builds on7
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Prompting Large Language Model for Machine Translation: A Case StudyBiao Zhang, Barry Haddow, Alexandra BirchICML 2023 · 402 citations
- Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A ReviewFred Philippy, Siwen Guo, Shohreh HaddadanACL 2023 · 9 citations
- SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian LanguagesHoly Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James V. Miranda et al.EMNLP 2024 · 8 citations
- To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language ProcessingSireesh Gururaja, Amanda Bertsch, Clara Na, David Gray Widder et al.EMNLP 2023 · 6 citations
Related papers
- Batayan: A Filipino NLP benchmark for evaluating Large Language ModelsJann Railey Montalan, Jimson Paulo Layacan, David Demitri Africa, Richell Isaiah Flores et al.ACL 2025 · 4 citations
- VMLU Benchmarks: A comprehensive benchmark toolkit for Vietnamese LLMsCuc Thi Bui, Nguyen Truong Son, Trang Van Truong, Viet Lam Phung et al.ACL 2025
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou et al.EMNLP 2024 · 59 citations
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
- LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian LanguagesAlham Fikri Aji, Trevor CohnEMNLP 2025 · 2 citations
