Holistically Evaluating the Environmental Impact of Creating Language Models
Jacob Morrison, Clara Na, Jared Fernandez, Tim Dettmers, Emma Strubell, Jesse Dodge
Abstract
As the performance of artificial intelligence systems has dramatically increased, so too has the environmental impact of creating these systems. While many model developers release estimates of the power consumption and carbon emissions from the final training runs for their latest models, there is comparatively little transparency into the impact of model development, hardware manufacturing, and total water usage throughout. In this work, we estimate the real-world environmental impact of developing a series of language models, ranging from 20 million to 13 billion active parameters, trained on up to 5.6 trillion tokens each. When accounting for hardware manufacturing, model development, and our final training runs, we find that our series of models released 493 metric tons of carbon emissions, equivalent to powering about 98 homes in the United States for one year, and consumed 2.769 million liters of water, equivalent to about 24.5 years of water usage by a person in the United States, even though our data center is extremely water-efficient. We measure and report the environmental impact of our model development; to the best of our knowledge we are the first to do so for LLMs, and we find that model development, the impact of which is generally not disclosed by most model developers, amounted to 50% of that of training. By looking at detailed time series data for power consumption, we also find that power usage throughout training is not consistent, fluctuating between 15% and 85% of our hardware's maximum power draw, with negative implications for grid-scale planning as demand continues to grow. We close with a discussion on the continued difficulty of estimating the environmental impact of AI systems, and key takeaways for model developers and the public at large.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 363dcc47-e905-435e-b4dc-abd19d4ee1cfCited by top-tier papers5
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth ObservationHenry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng et al.CVPR 2026 · 28 citations
- Predicting LLM Reasoning Performance with Small Proxy ModelWoosung Koh, Juyoung Suk, Sungjun Han, Se-Young Yun et al.ICLR 2026 · 5 citations
- Environmental Footprints of Query Processing: A Vision for Sustainable Database ArchitecturesMichail Bachras, Hans-Arno JacobsenVLDB 2025 · 1 citation
- Energy Considerations of Large Language Model Inference and Efficiency OptimizationsJared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk et al.ACL 2025
- Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at ScaleXinlei Wang, Ruibo Ming, Jing Qiu, Junhua Zhao et al.ICML 2026
Builds on1
Related papers
- LLMCarbon: Modeling the End-to-End Carbon Footprint of Large Language ModelsAhmad Faiz, Sotaro Kaneda, Ruhan Wang, Rita Chukwunyere Osi et al.ICLR 2024 · 129 citations
- Towards Climate Awareness in NLP ResearchDaniel Hershcovich, Nicolas Webersinke, Mathias Kraus, Julia Anna Bingler et al.EMNLP 2022 · 32 citations
- Unveiling the Uncertainty in Embodied and Operational Carbon of Large AI Models through a Probabilistic Carbon Accounting ModelXiaoyang Zhang, Fang He, Yang Deng, Dan WangNeurIPS 2025 · 3 citations
- TokenPowerBench: Benchmarking the Power Consumption of LLM InferenceChenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao et al.AAAI 2026 · 12 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
