Hugging Carbon: Quantifying the Training Carbon Emissions of AI Models at Scale
Xinlei Wang, Ruibo Ming, Jing Qiu, Junhua Zhao, Jinjin Gu
摘要
The scaling-law era has transformed artificial intelligence (AI) from research into a global industry, but its rapid growth also raises concerns over energy usage, carbon emissions, and environmental sustainability. Unlike traditional sectors, the AI industry still lacks systematic carbon accounting methods that support large-scale estimates without reproducing the original training process. This leaves open questions about how large the problem is today and how large it might be in the near future. Given its central role in hosting open-source AI models, the Hugging Face (HF) platform provides a large-scale and publicly accessible corpus for carbon accounting. We estimate aggregate training emissions of HF opensource models using available emissions, energy, compute, and model metadata. To address uneven disclosure quality, we introduce a tiered approach to handle incomplete metadata, supported by empirical regressions that assess estimation reliability. We further introduce AI training carbon intensity (ATCI, emissions per compute), a metric to assess the sustainability efficiency of model training. Our results show that training the most popular open-source models (with over 5,000 downloads) has already resulted in approximately 6.0 × 10 4 metric tons of carbon emissions. Overall, this paper provides a scalable, empirically grounded framework for estimating training emissions from incomplete disclosures and informing future carbon reporting standards in the AI industry. Data and code are available at https://github.com/insait-insti tute/HuggingCarbon .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang 等USENIX ATC 2023 · 被引用 191 次
- Efficient Large Scale Language Modeling with Mixtures of ExpertsMikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov 等EMNLP 2022 · 被引用 71 次
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 被引用 33 次
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong 等EuroSys 2026 · 被引用 5 次
- Holistically Evaluating the Environmental Impact of Creating Language ModelsJacob Morrison, Clara Na, Jared Fernandez, Tim Dettmers 等ICLR 2025
相关 Paper
- Unveiling the Uncertainty in Embodied and Operational Carbon of Large AI Models through a Probabilistic Carbon Accounting ModelXiaoyang Zhang, Fang He, Yang Deng, Dan WangNeurIPS 2025 · 被引用 3 次
- Your Space is My Zone: Demystifying the Security Risks of AI-Powered Applications on Pre-Trained Model HubsYacong Gu, Lingyun Ying, Zidong Zhang, Yingyuan Pu 等CCS 2026 · 被引用 1 次
- LLMCarbon: Modeling the End-to-End Carbon Footprint of Large Language ModelsAhmad Faiz, Sotaro Kaneda, Ruhan Wang, Rita Chukwunyere Osi 等ICLR 2024 · 被引用 129 次
- Green AI: Do Deep Learning Frameworks Have Different Costs?Stefanos Georgiou, Maria Kechagia, Tushar Sharma, Federica Sarro 等ICSE 2022 · 被引用 90 次
- The Hidden Joules: Evaluating the Energy Consumption of Vision Backbones for Progress Towards More Efficient Model InferenceZeyu Yang, Wesley ArmourICML 2025
