MGAL: A Multilingual Granularity-Aware Long-Context Benchmark
Chunhan Li, Chenglin Xu, Zongyang Zhang, Jiale Liu, Zhuoxi Rao, Xudong jia, JUNXIU HE, Menglin Yang, Wenjuan Gong, Zhengzhe Liu, Chengwei Qin
摘要
Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the first multilingual, granularity- and position-aware long-context benchmark. MGAL is constructed from United Nations (UN) reports spanning 8K to 128K tokens across the six official UN languages. It covers four coherent levels of linguistic granularity (word, sentence, paragraph, and document) and further stratifies entries by their position within the document (begin, middle, and end), indexed at both the document and paragraph levels. This design enables systematic diagnosis of multilingual long-context comprehension across different granularities. Through extensive experiments and analyses, we find that: (1) LLMs perform well at word-level tasks but struggle with coarser-grained ones; and (2) Closed-source models retain a clear performance advantage in lower-resource languages. We further identify two new challenges: (1) Under local semantic crowding, where neighboring sentences share topics and entities, models tend to follow surface cues (e.g., connectives like 'however' or repeated entities) rather than the discourse role of the sentence in surrounding context (e.g., background, outcome); and (2) A gap between fluency and consistency in generated outputs, where models produce text that reads smoothly but drifts from the source facts. In addition, we observe several patterns in line with prior studies, including reliance on nearby evidence and reuse of options under uncertainty.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
- M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun 等ACL 2024 · 被引用 3 次
- Squeezed Attention: Accelerating Long Context Length LLM InferenceColeman Richard Charles Hooper, Sehoon Kim, Hiva Mohammadzadeh, Monishwaran Maheswaran 等ACL 2025
相关 Paper
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 等ICLR 2023 · 被引用 52 次
- Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng 等ACL 2025 · 被引用 3 次
- Extending Automatic Machine Translation Evaluation to Book-Length DocumentsKuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh 等EMNLP 2025
- NovelQA: Benchmarking Question Answering on Documents Exceeding 200K TokensCunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu 等ICLR 2025
- INCLUDE: Evaluating Multilingual Language Understanding with Regional KnowledgeAngelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen 等ICLR 2025
