UniRTL: Unifying Code and Graph for Robust RTL Representation Learning
Yi Liu, Hongji Zhang, Lei Chen, Mingxuan Yuan, Qiang Xu
Abstract
Developing effective representations for register transfer level (RTL) designs is crucial for accelerating the hardware design workflow. Existing approaches, however, typically rely on a single data modality, either the RTL code or its associated graph-based representation, limiting the expressiveness and generalization ability of the learned representations. For RTL, the control data flow graph (CDFG) offers a comprehensive structural representation that preserves complete information, while the code modality explicitly encodes semantic and functional information. We argue that integrating these complementary modalities is essential for a thorough understanding of RTL designs. To this end, we propose UniRTL, a multimodal pretraining framework that learns unified RTL representations by jointly leveraging code and CDFG. UniRTL achieves fine-grained alignment between code and graph through mutual masked modeling and employs a hierarchical training strategy that incorporates a pretrained graph-aware tokenizer and staged alignment of text ( i.e. , functional summary) and code prior to graph integration. We evaluate UniRTL on two downstream tasks, performance prediction and code retrieval, under multiple settings. Experimental results show that UniRTL consistently outperforms prior methods, establishing it as a more robust and powerful foundation for advancing hardware design automation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 993fce70-a962-48b6-b1e4-3a3d58094e9dBuilds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
- BetterV: Controlled Verilog Generation with Discriminative GuidanceZehua Pei, Hui-Ling Zhen, Mingxuan Yuan, Yu Huang et al.ICML 2024 · 155 citations
Related papers
- Topology Matters in RTL Circuit Representation LearningMingyu Zhao, Xun He, Jiawei Liu, Jianwang Zhai et al.ICLR 2026
- Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph LearningYi Liu, Hongji Zhang, Yiwen Wang, Dimitrios Tsaras et al.ICML 2026 · 1 citation
- DynamicRTL: RTL Representation Learning for Dynamic Circuit BehaviorRuiyang Ma, Yunhao Zhou, Yipeng Wang, Yi Liu et al.AAAI 2026
- FAIR: Flow Type-Aware Pre-Training of Compiler Intermediate RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, David Lo et al.ICSE 2024 · 5 citations
- NetTAG: A Multimodal RTL-and-Layout-Aligned Netlist Foundation Model via Text-Attributed GraphWenji Fang, Wenkai Li, Shang Liu, Yao Lu et al.DAC 2025 · 10 citations
