NLP Libraries, Energy Consumption and Runtime: An Empirical Study
Rajrupa Chattaraj, Sridhar Chimalakonda
Abstract
In the realm of natural language processing (NLP), the rising computational demands of modern models bring energy efficiency to the forefront of sustainable computing. Preprocessing tasks, such as tokenization, stemming, and POS tagging, are critical steps in transforming raw text into structured formats suitable for machine learning models. However, despite their widespread use in numerous NLP pipelines, little attention has been given to their energy consumption. This empirical study evaluates and compares the energy consumption and runtime performance of three popular NLP libraries— NLTK , spaCy , and Gensim —across six common preprocessing tasks. We conducted a comprehensive comparison using three distinct datasets and six preprocessing tasks. Energy consumption was measured using the Intel-RAPL and NVIDIA-SMI interfaces, while runtime performance was recorded across all library-task combinations. The results reveal substantial discrepancies in energy consumption across the three libraries, with up to 93% of cases exhibiting significant variations. Gensim showed superior efficiency in tokenization and stemming, while spaCy excelled in tasks like POS tagging and Named Entity Recognition (NER). These findings underscore the potential for optimizing NLP preprocessing tasks for energy efficiency. Our study highlights the untapped potential for improving energy efficiency in NLP pipelines. These insights emphasize the need for more focused research into energy-efficient NLP techniques, especially in the preprocessing phase, to support the development of greener, more sustainable computational models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Green AI: Do Deep Learning Frameworks Have Different Costs?Stefanos Georgiou, Maria Kechagia, Tushar Sharma, Federica Sarro et al.ICSE 2022 · 90 citations
- Energy Considerations of Large Language Model Inference and Efficiency OptimizationsJared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk et al.ACL 2025
- TokenPowerBench: Benchmarking the Power Consumption of LLM InferenceChenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao et al.AAAI 2026 · 12 citations
- IrEne: Interpretable Energy Prediction for TransformersQingqing Cao, Yash Kumar Lal, Harsh Trivedi, Aruna Balasubramanian et al.ACL 2021
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 96 citations
