CoFrGeNet: Continued Fraction Architectures for Language Generation
Amit Dhurandhar, Vijil Chenthamarakshan, Dennis Wei, Tejaswini Pedapati, Karthikeyan Natesan Ramamurthy, Rahul Nair
摘要
Transformers are arguably the preferred architecture for language generation. In this paper, inspired by continued fractions, we introduce a new function class for generative modeling. The architecture family implementing this function class is named CoFrGeNets -Continued Fraction Generative Networks. We design novel architectural components based on this function class that can replace Multihead Attention and Feed-Forward Networks in Transformer blocks while requiring much fewer parameters. We derive custom gradient formulations to optimize the proposed components more accurately and efficiently than using standard PyTorch-based gradients. Our components are a plug-in replacement requiring little change in training or inference procedures that have already been put in place for Transformer-based models thus making our approach easy to incorporate in large industrial workflows. We experiment on two very different transformer architectures GPT2-xl (1.5B) and Llama3 (3.2B), where the former we pre-train on OpenWebText and GneissWeb, while the latter we pre-train on the docling data mix which consists of nine different datasets. Results show that the performance on downstream classification, Q& A, reasoning and text understanding tasks of our models is competitive and sometimes even superior to the original models with 2 3 to 1 2 the parameters and shorter pre-training time. We believe that future implementations customized to hardware will further bring out the true potential of our architectures. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
相关 Paper
- CoFrNets: Interpretable Neural Architecture Inspired by Continued FractionsIsha Puri, Amit Dhurandhar, Tejaswini Pedapati, Karthikeyan Shanmugam 等NeurIPS 2021 · 被引用 13 次
- Language Models Are Implicitly ContinuousSamuele Marro, Davide Evangelista, Xuanqiang Angelo Huang, Emanuele La Malfa 等ICLR 2025
- Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsShansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye 等ICLR 2025
- Flextron: Many-in-One Flexible Large Language ModelRuisi Cai, Saurav Muralidharan, Greg Heinrich, Hongxu Yin 等ICML 2024 · 被引用 38 次
- Mini-Sequence Transformers: Optimizing Intermediate Memory for Long Sequences TrainingCheng Luo, Jiawei Zhao, Zhuoming Chen, Beidi Chen 等NeurIPS 2024 · 被引用 6 次
