Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at Scale
Xinlei Wang, Ruibo Ming, Jing Qiu, Junhua Zhao, Jinjin Gu
Abstract
The scaling-law era has transformed artificial intelligence (AI) from research into a global industry, but its rapid growth also raises concerns over energy usage, carbon emissions, and environmental sustainability. Unlike traditional sectors, the AI industry still lacks systematic carbon accounting methods that support large-scale estimates without reproducing the original training process. This leaves open questions about how large the problem is today and how large it might be in the near future. Given its central role in hosting open-source AI models, the Hugging Face (HF) platform provides a large-scale and publicly accessible corpus for carbon accounting. We estimate aggregate training emissions of HF opensource models using available emissions, energy, compute, and model metadata. To address uneven disclosure quality, we introduce a tiered approach to handle incomplete metadata, supported by empirical regressions that assess estimation reliability. We further introduce AI training carbon intensity (ATCI, emissions per compute), a metric to assess the sustainability efficiency of model training. Our results show that training the most popular open-source models (with over 5,000 downloads) has already resulted in approximately 6.0 × 10 4 metric tons of carbon emissions. Overall, this paper provides a scalable, empirically grounded framework for estimating training emissions from incomplete disclosures and informing future carbon reporting standards in the AI industry. Data and code are available at https://github.com/insait-insti tute/HuggingCarbon .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a672f073-1fb4-451d-97db-cc80a1dfd443Builds on6
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang et al.USENIX ATC 2023 · 191 citations
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov et al.EMNLP 2022 · 71 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong et al.EuroSys 2026 · 5 citations
- Holistically Evaluating the Environmental Impact of Creating Language ModelsJacob Morrison, Clara Na, Jared Fernandez, Tim Dettmers et al.ICLR 2025
Related papers
- Unveiling the Uncertainty in Embodied and Operational Carbon of Large AI Models through a Probabilistic Carbon Accounting ModelXiaoyang Zhang, Fang He, Yang Deng, Dan WangNeurIPS 2025 · 3 citations
- Your Space is My Zone: Demystifying the Security Risks of AI-Powered Applications on Pre-Trained Model HubsYacong Gu, Lingyun Ying, Zidong Zhang, Yingyuan Pu et al.CCS 2026 · 1 citation
- LLMCarbon: Modeling the End-to-End Carbon Footprint of Large Language ModelsAhmad Faiz, Sotaro Kaneda, Ruhan Wang, Rita Chukwunyere Osi et al.ICLR 2024 · 129 citations
- Green AI: Do Deep Learning Frameworks Have Different Costs?Stefanos Georgiou, Maria Kechagia, Tushar Sharma, Federica Sarro et al.ICSE 2022 · 90 citations
- The Hidden Joules: Evaluating the Energy Consumption of Vision Backbones for Progress Towards More Efficient Model InferenceZeyu Yang, Wesley ArmourICML 2025
