FilBench: Can LLMs Understand and Generate Filipino?
Lester James Validad Miranda, Elyanah Aco, Conner G. Manuel, Jan Christian Blaise Cruz, Joseph Marvin Imperial
摘要
Despite the impressive performance of LLMs on English-based tasks, little is known about their capabilities in specific languages such as Filipino. In this work, we address this gap by introducing FILBENCH, a Filipino-centric benchmark designed to evaluate LLMs across a diverse set of tasks and capabilities in Filipino, Tagalog, and Cebuano. We carefully curate the tasks in FILBENCH to reflect the priorities and trends of NLP research in the Philippines such as Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. By evaluating 27 state-of-the-art LLMs on FILBENCH, we find that several LLMs suffer from reading comprehension and translation capabilities. Our results indicate that FILBENCH is challenging, with the best model, GPT-4o, achieving only a score of 72.23%. Moreover, we also find that models trained specifically for Southeast Asian languages tend to underperform on FIL-BENCH, with the highest-performing model, SEA-LION v3 70B, achieving only a score of 61.07%. Our work demonstrates the value of curating language-specific LLM benchmarks to aid in driving progress on Filipino NLP and increasing the inclusion of Philippine languages in LLM development.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Prompting Large Language Model for Machine Translation: A Case StudyBiao Zhang, Barry Haddow, Alexandra BirchICML 2023 · 被引用 402 次
- Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A ReviewFred Philippy, Siwen Guo, Shohreh HaddadanACL 2023 · 被引用 9 次
- SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian LanguagesHoly Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James V. Miranda 等EMNLP 2024 · 被引用 8 次
- To Build Our Future, We Must Know Our Past: Contextualizing Paradigm Shifts in Natural Language ProcessingSireesh Gururaja, Amanda Bertsch, Clara Na, David Gray Widder 等EMNLP 2023 · 被引用 6 次
相关 Paper
- Batayan: A Filipino NLP benchmark for evaluating Large Language ModelsJann Railey Montalan, Jimson Paulo Layacan, David Demitri Africa, Richell Isaiah Flores 等ACL 2025 · 被引用 4 次
- VMLU Benchmarks: A comprehensive benchmark toolkit for Vietnamese LLMsCuc Thi Bui, Nguyen Truong Son, Trang Van Truong, Viet Lam Phung 等ACL 2025
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou 等EMNLP 2024 · 被引用 59 次
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun 等ACL 2024
- LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian LanguagesAlham Fikri Aji, Trevor CohnEMNLP 2025 · 被引用 2 次
