CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
Irene Wang, Mostafa Elhoushi, Ekin Sumbul, Samuel Hsia, Daniel Jiang, Newsha Ardalani, Divya Mahajan, Carole-Jean Wu, Bilge Acun
Abstract
Machine learning solutions are rapidly adopted to enable a variety of key use cases, from conversational AI assistants to scientific discovery. This growing adoption is expected to increase the associated lifecycle carbon footprint, including both operational carbon from training and inference and embodied carbon from AI hardware manufacturing. We introduce -- the first carbon-aware co-optimization framework for Transformer-based models and hardware accelerators. By integrating both operational and embodied carbon into early-stage design space exploration, enables sustainability-driven model architecture and hardware accelerator co-design that reveals fundamentally different trade-offs than latency- or energy-centric approaches. Evaluated across a range of Transformer models, consistently demonstrates the potential to reduce total carbon emissions -- by up to 30% -- while maintaining accuracy and latency. We further highlight its extensibility through a focused case study on multi-modal models. Our results emphasize the need for holistic optimization methods that prioritize carbon efficiency without compromising model capability and execution time performance. The source code of is available at https://github.com/facebookresearch/CATransformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5b17153-dfd1-4628-98fb-01db9c571266Cited by top-tier papers1
Ask how each one uses itBuilds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Parallel Bayesian Optimization of Multiple Noisy Objectives with Expected Hypervolume ImprovementSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2021 · 276 citations
- HAT: Hardware-Aware Transformers for Efficient Natural Language ProcessingHanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai et al.ACL 2020 · 215 citations
- TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight InheritanceKan Wu, Houwen Peng, Zhenghong Zhou, Bin Xiao et al.ICCV 2023 · 118 citations
Related papers
- ACT: designing sustainable computer systems with an architectural carbon modeling toolUdit Gupta, Mariam Elgamal, Gage Hills, Gu-Yeon Wei et al.ISCA 2022 · 176 citations
- LLMCarbon: Modeling the End-to-End Carbon Footprint of Large Language ModelsAhmad Faiz, Sotaro Kaneda, Ruhan Wang, Rita Chukwunyere Osi et al.ICLR 2024 · 129 citations
- CE-NAS: An End-to-End Carbon-Efficient Neural Architecture Search FrameworkYiyang Zhao, Yunzhuo Liu, Bo Jiang, Tian GuoNeurIPS 2024 · 10 citations
- FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation PredictionYubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao et al.ISCA 2023 · 113 citations
- Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit ViewYanran Wu, Inez Hua, Yi DingACL 2025 · 15 citations
