SYNC: A Synthetic Long-Context Understanding Benchmark for Controlled Comparisons of Model Capabilities
Shuyang Cao, Kaijian Zou, Lu Wang
Abstract
Recently, researchers have turned to synthetic tasks for evaluating long-context capabilities of large language models (LLMs) , as they offer more flexibility than realistic benchmarks in scaling both input length and dataset size. However, existing synthetic tasks typically target narrow skill sets such as retrieving information from massive input, limiting their ability to comprehensively assess model capabilities. Furthermore, existing benchmarks often pair each task with a different input context, creating confounding factors that prevent fair crosstask comparison. To address these limitations, we introduce SYNC, a new evaluation suite of synthetic tasks spanning domains including graph understanding and translation. Each domain includes three tasks designed to test a wide range of capabilities-from retrieval, to multi-hop tracking, and to global context understanding that that requires chain-of-thought (CoT) reasoning. Crucially, all tasks share the same context, enabling controlled comparisons of model performance. We evaluate 14 LLMs on SYNC and observe substantial performance drops on more challenging tasks, underscoring the benchmark's difficulty. Additional experiments highlight the necessity of CoT reasoning and demonstrate that SYNC poses a robust challenge for future models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- SQuALITY: Building a Long-Document Summarization Dataset the Hard WayAlex Wang, Richard Yuanzhe Pang, Angelica Chen, Jason Phang et al.EMNLP 2022 · 19 citations
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAMinzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao et al.EMNLP 2024 · 9 citations
Related papers
- M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun et al.ACL 2024 · 3 citations
- LongCoT: Benchmarking Long-Horizon Chain-of-Thought ReasoningSumeet Motwani, Daniel Nichols, Charles London, Peggy Li et al.ICML 2026 · 2 citations
- GSM-∞: How Do your LLMs Behave over Infinitely Increasing Reasoning Complexity and Context Length?Yang Zhou, Hongyi Liu, Zhuoming Chen, Yuandong Tian et al.ICML 2025
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
- HELMET: How to Evaluate Long-context Models Effectively and ThoroughlyHoward Yen, Tianyu Gao, Minmin Hou, Ke Ding et al.ICLR 2025
