Lune

ISCA2026顶会

TUSQ: Tracking, Uncomputation, and Sampling for Noisy Quantum Simulation

Siddharth Dangwal, Tina Oberoi, Ajay Sailopal, Dhirpal Shah, Frederic T. Chong

2026年份

摘要

Quantum computers have improved in size and quality in recent years, enabling the execution of complex circuits. However, for most researchers, access to compute time is limited. This necessitates the development of simulators that mimic noisy quantum hardware accurately and scalably. The ideal way to simulate noisy systems is via Density Matrix Simulation (DMS). However, its high memory footprint limits its scalability. Consequently, noisy simulations are performed in two steps: (a) sampling multiple circuits with fixed noisy gates from the stochastic noise channels, (b) performing the State Vector Simulations (SVS) of these circuits and averaging their output to obtain the effective noisy simulation result. This often leads to a substantial increase in compute overhead, slowing down the simulation. Existing methods solve this problem by caching critical intermediate results in memory and reusing them. However, when a simulation task is both compute and memory-intensive, we need to eliminate computational overheads without incurring extra memory overheads. To enable fast simulation in the compute and memory bound regime, we propose TUSQ - Tracking, Uncomputation, and Sampling for Noisy Quantum Simulation. TUSQ is composed of two modules: the Error Characterization Module (ECM), and Depth First Tree Traversal (DFTT). The ECM characterizes errors so that the simulator can eliminate redundant circuit instances (via ER Tallying and ER Commutation), followed by importance sampling (in the Pruning stage), significantly reducing the number of circuits to be simulated relative to the baseline strategy of simulating all circuits. This is followed by DFTT, which computes the statevectors for these sampled circuits efficiently by taking advantage of circuit similarity, representing similar circuits in a tree and using computation and uncomputation to traverse the tree efficiently. TUSQ is evaluated for a total of 198 benchmarks, executed for 1 million shots and reports an average speedup of 59.06×59.06 \times and 13.38×13.38 \times over Qiskit and CUDA-Q, with a maximum speedup of 7878.03×7878.03 \times and 439.38×439.38 \times respectively. We also compare TUSQ against TQSim in the time and memory critical regime. We observe an average and maximum speedup of 39.32×39.32 \times and 3134.31×, respectively.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖