CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
Ahmed Heakl, Gustavo Bertolo Stahl, Sarim Hashmi, Seung Hun Eddie Han, Mukul Ranjan, Arina Kharlamova, Salman Khan, Abdulrahman Mahmoud
Abstract
Cross-architecture GPU code transpilation is essential for unlocking low-level hardware portability, yet no scalable solution exists. We introduce CASS, the first dataset and model suite for source-and assembly-level GPU translation (CUDA ↔ HIP, SASS ↔ RDNA3). CASS contains 60k verified host-device code pairs, enabling learning-based translation across both ISA and runtime boundaries. We generate each sample using our automated pipeline that scrapes, translates, compiles, and aligns GPU programs across vendor stacks. Leveraging CASS, we train a suite of domainspecific translation models that achieve 88.2% accuracy on CUDA→HIP and 69.1% on SASS→RDNA3, outperforming commercial baselines including GPT-5.1, Claude-4.5, and Hipify by wide margins. Generated code matches native performance in 95% of cases, preserving both runtime and memory behavior. To support rigorous evaluation, we introduce CASS-Bench, a curated benchmark spanning 18 GPU domains with ground-truth execution. All data, models, and evaluation tools will be released as open source to support progress in GPU compiler tooling, binary compatibility, and LLM-guided code translation. * † Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b79d8fc-dd45-452a-a664-0671efd887a9Builds on1
Related papers
- Task-Based Tensor Computations on Modern GPUsRohan Yadav, Michael Garland, Alex Aiken, Michael BauerPLDI 2025 · 5 citations
- BabelTower: Learning to Auto-parallelized Program TranslationYuanbo Wen, Qi Guo, Qiang Fu, Xiaqing Li et al.ICML 2022 · 27 citations
- Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code TranslationLe Chen, Nuo Xu, Winson Chen, Bin Lei et al.ACL 2026 · 6 citations
- The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSLBo Qiao, M. Akif Özkan, Jürgen Teich, Frank HannigDAC 2020 · 8 citations
- High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel ConstructsWilliam S. Moses, Ivan R. Ivanov, Jens Domke, Toshio Endo et al.PPoPP 2023 · 27 citations
