A Foundation Model for Error Correction Codes
Yoni Choukroun, Lior Wolf
Abstract
In recent years, Artificial Intelligence has undergone a paradigm shift with the rise of foundation models, which are trained on large amounts of data, typically in a self-supervised way, and can then be adapted to a wide range of downstream tasks. In this work, we propose the first foundation model for Error Correction Codes. This model is trained on multiple codes and can then be applied to an unseen code. To enable this, we extend the Transformer architecture in multiple ways: (1) a code-invariant initial embedding, which is also position-and lengthinvariant, (2) a learned modulation of the attention maps that is conditioned on the Tanner graph, and (3) a length-invariant code-aware noise prediction module that is based on the parity-check matrix. The proposed architecture is trained on multiple short-and medium-length codes and is able to generalize to unseen codes. Its performance on these codes matches and even outperforms the state of the art, despite having a smaller capacity than the leading code-specific transformers. The suggested framework therefore demonstrates, for the first time, the benefits of learning a universal decoder rather than a decoder optimized for a given code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55a65aa4-80be-4af1-a773-17f4b8be8360Cited by top-tier papers5
- Efficient Message-Passing Transformer for Error Correcting CodesSeong-Joon Park, Taewoo Park, Hee-Youl Kwak, Sang-Hyo Kim et al.ICLR 2026 · 29 citations
- Learning Linear Block Error Correction CodesYoni Choukroun, Lior WolfICML 2024 · 18 citations
- Drop-in Circulant Structural Priors for Transformer Decoding of Cyclic CodesShuai Xiao, Weijun Fang, Qiaosheng ZhangICML 2026
- CrossMPT: Cross-attention Message-passing Transformer for Error Correcting CodesSeong-Joon Park, Heeyoul Kwak, Sang-Hyo Kim, Yongjune Kim et al.ICLR 2025
- Score Based Error Correcting Code DecoderAlon Helvits, Eliya NachmaniICML 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Error Correction Code TransformerYoni Choukroun, Lior WolfNeurIPS 2022 · 121 citations
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin et al.NeurIPS 2024 · 14 citations
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- CBraMod: A Criss-Cross Brain Foundation Model for EEG DecodingJiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou et al.ICLR 2025
- Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning TasksXingcheng Xu, Zibo Zhao, Haipeng Zhang, Yanqing YangACL 2025 · 2 citations
