XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition
Zhiren Gong, Tiantong Wu, Jiaming Zhang, Fuyao Zhang, CHE WANG, Yurong Hao, Yikun Hou, Foo Ping, Yilei Zhao, Fei Huang, Chau Yuen, Wei Yang Bryan Lim
Abstract
Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on singleturn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDOMAINBENCH, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stresstesting from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domainmixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse. We have code in GitHub repository, project page in XDomainBench, and dataset in Hugging Face.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b8d1082-6181-4530-8ef3-24796c24a7f5Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
Related papers
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code GenerationQiaosheng Chen, Yang Liu, Lei Li, Kai Chen et al.ICML 2026 · 1 citation
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language ModelsZhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia et al.ICLR 2026 · 24 citations
- Towards Multimodal Data-Driven Scientific Discovery Powered by LLM AgentsFan Liu, Xiaozhao Zeng, Hao LiuICLR 2026
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMsMohammad Ramezanali, Mo Vazifeh, Paolo SantiEMNLP 2025
