The Framework Tax: Disparities Between Inference Efficiency in NLP Research and Deployment
Jared Fernandez, Jacob Kahn, Clara Na, Yonatan Bisk, Emma Strubell
Abstract
Increased focus on the computational efficiency of systems in natural language processing has motivated the design of efficient model architectures and improvements to underlying hardware accelerators. However, the resulting increases in computational throughput and reductions in floating point operations have not directly translated to improvements in wall-clock inference latency. We demonstrate that these discrepancies can be largely attributed to bottlenecks introduced by deep learning frameworks. We denote this phenomena as the framework tax, and observe that the disparity is growing as hardware speed increases over time. In this work, we examine this phenomena through a series of case studies analyzing the effects of model design decisions, framework paradigms, and hardware platforms on total model latency. Based on our findings, we provide actionable recommendations to researchers and practitioners aimed at narrowing the gap between efficient NLP model research and practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b73cdbb-4f55-4756-82a2-973fee0e9effCited by top-tier papers2
- Energy Considerations of Large Language Model Inference and Efficiency OptimizationsJared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk et al.ACL 2025
- Optimized Speculative Sampling for GPU Hardware AcceleratorsDominik Wagner, Seanie Lee, Ilja Baumann, Philipp Seeberger et al.EMNLP 2024
Builds on16
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
Related papers
- Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer ModelsDeepak Narayanan, Keshav Santhanam, Peter Henderson, Rishi Bommasani et al.NeurIPS 2023 · 14 citations
- Explainable-DSE: An Agile and Explainable Exploration of Efficient HW/SW Codesigns of Deep Learning Accelerators Using Bottleneck AnalysisShail Dave, Tony Nowatzki, Aviral ShrivastavaASPLOS 2023 · 7 citations
- A full-stack search technique for domain optimized deep learning acceleratorsDan Zhang, Safeen Huda, Ebrahim M. Songhori, Kartik Prabhu et al.ASPLOS 2022 · 48 citations
- Vector-Vector-Matrix Architecture: A Novel Hardware-Aware Framework for Low-Latency Inference in NLP ApplicationsMatthew Khoury, Rumen Dangovski, Longwu Ou, Preslav Nakov et al.EMNLP 2020 · 2 citations
- Understanding performance problems in deep learning systemsJunming Cao, Bihuan Chen, Chao Sun, Longjie Hu et al.FSE 2022 · 33 citations
