ZipZap: Efficient Training of Language Models for Large-Scale Fraud Detection on Blockchain
Sihao Hu, Tiansheng Huang, Ka-Ho Chow, Wenqi Wei, Yanzhao Wu, Ling Liu
Abstract
Language models (LMs) have demonstrated superior performance in detecting fraudulent activities on Blockchains. Nonetheless, the sheer volume of Blockchain data results in excessive memory and computational costs when training LMs from scratch, limiting their capabilities to large-scale applications. In this paper, we present ZipZap, a framework tailored to achieve both parameter and computational efficiency when training LMs on large-scale transaction data. First, with the frequency-aware compression, an LM can be compressed down to a mere 7.5% of its initial size with an imperceptible performance dip. This technique correlates the embedding dimension of an address with its occurrence frequency in the dataset, motivated by the observation that embeddings of low-frequency addresses are insufficiently trained and thus negating the need for a uniformly large dimension for knowledge representation. Second, ZipZap accelerates the speed through the asymmetric training paradigm: It performs transaction dropping and cross-layer parameter-sharing to expedite the pre-training process, while revert to the standard training paradigm for fine-tuning to strike a balance between efficiency and efficacy, motivated by the observation that the optimization goals of pre-training and fine-tuning are inconsistent. Evaluations on real-world, large-scale datasets demonstrate that ZipZap delivers notable parameter and computational efficiency improvements for training LMs. Our implementation is available at: https://github.com/git-disl/ZipZap.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0df168cb-acdb-41ae-b4db-de0a1ce3cdc9Cited by top-tier papers6
- Know Your Account: Double Graph Inference-Based Account De-Anonymization on EthereumShuyi Miao, Wangjie Qiu, Hongwei Zheng, Qinnan Zhang et al.ICDE 2025 · 3 citations
- HiLoMix: Robust High- and Low-Frequency Graph Learning Framework for Mixing Address AssociationXiaofan Tu, Tiantian Duan, Shuyi Miao, Hanwen Zhang et al.AAAI 2026 · 1 citation
- Evolving Proxy Kills Drift: Data-Efficient Streaming Time Series Anomaly DetectionQing Wei, Hao Miao, Yan Zhao, Kai Zheng et al.WWW 2026 · 1 citation
- SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement LearningPeidong Wang, Zhiming Ma, Xin Dai, Yongkang Liu et al.ACL 2026 · 1 citation
- IGT4ETH: An Isotropic Pre-trained Graph Transformer for Ethereum Account ClassificationAo Liu, Yanmei Zhang, Youwei Wang, Qiang DuanAAAI 2026
Related papers
- BlockScan: Detecting Anomalies in Blockchain TransactionsJiahao Yu, Xian Wu, Hao Liu, Wenbo Guo et al.NeurIPS 2025 · 7 citations
- zip2zip: Inference-Time Adaptive Tokenization via Online CompressionSaibo Geng, Nathan Ranchin, Yunzhen Yao, Maxime Peyrard et al.NeurIPS 2025 · 5 citations
- BERT4ETH: A Pre-trained Transformer for Ethereum Fraud DetectionSihao Hu, Zhen Zhang, Bingqiao Luo, Shengliang Lu et al.WWW 2023 · 97 citations
- EarlyBERT: Efficient BERT Training via Early-bird Lottery TicketsXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan et al.ACL 2021
- Phishing in Wonderland: Evaluating Learning-Based Ethereum Phishing Transaction Detection and PitfallsAhod Alghuried, David MohaisenNDSS 2026 · 4 citations
