QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
Shvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand, Varun Gohil, Jeffrey Ma, Jason Yik, Zishen Wan, Jessica A. Quaye, Elisavet Alvanaki, Avinash Kumar, Chandrashis Mazumdar
Abstract
The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) evaluations. To this end, we present QuArch (pronounced ‘quark’), the first benchmark designed to facilitate the development and evaluation of LLM knowledge and reasoning capabilities specifically in computer architecture. QuArch v1.0 provides a comprehensive collection of 2,671 expert-validated question-answer (QA) pairs covering various aspects of computer architecture, including processor design, memory systems, and interconnection networks. Our evaluation reveals that while frontier models possess domain-specific knowledge, they struggle with skills that require higher-order thinking in computer architecture. Frontier model accuracies vary widely (from 34% to 73%) on these advanced questions, highlighting persistent gaps in architectural reasoning across analysis, design, and implementation QAs. Furthermore, via fine-tuning we find that QuArch can translate to improved performance on a realistic memory hierarchy design task, resulting in up to 1.99× more area-efficient solutions and up to 40% more viable solutions overall. By holistically assessing fundamental skills, QuArch provides a foundation for building and measuring LLM capabilities that can accelerate innovation in computing systems. The QuArch benchmark and leaderboard are publicly available at: https://quarch.ai/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ac59016-b934-42b7-844c-4a1fddda45ceBuilds on15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
- JEC-QA: A Legal-Domain Question Answering DatasetHaoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang et al.AAAI 2020 · 212 citations
Related papers
- PCB-Bench: Benchmarking LLMs for Printed Circuit Board Placement and RoutingJindong Li, Lianrong Chen, Bin Yang, Jiadong Zhu et al.ICLR 2026
- Can LLMs Reason Structurally? Benchmarking via the lens of Data StructuresYu He, Yingxi Li, Colin White, Ellen VitercikICML 2026 · 3 citations
- BIG-Bench Extra HardMehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch et al.ACL 2025
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIHuanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen et al.ICCV 2025 · 4 citations
- A²RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark GenerationQingchuan Ma, Yuexiao Ma, Yongkang Xie, Tianyu Xie et al.ICML 2026 · 1 citation
