article
Transformer models power modern natural language processing applications but incur substantial computational and energy costs in deployment. The inferencerelated costs of transformer models have received growing attention in recent years, particularly with respect to energy consumption and carbon footprint. However, few studies have attempted to benchmark these inference costs on a CPU and examine how model size influences these costs from a computational and environmental perspective. In this paper, we evaluate three pretrained transformer models-MiniLM (22M parameters), DistilBERT (66 M parameters), and BERT-base (110 M parameters)-under CPU-based inference. Using SST-2 sentences as a standardized NLP workload, we compare latency, throughput, energy consumption, and carbon footprint for these three models. Our results show that MiniLM achieved approximately <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$5.4 \times$</tex> lower latency, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$5.5 \times$</tex> higher throughput, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$4.9 \times$</tex> lower energy consumption, and <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$5.0 \times$</tex> lower carbon emissions than BERT-base. DistilBERT showed intermediate efficiency levels across all measured metrics, indicating that inference efficiency scales with model size. These results highlight the substantial computational and environmental advantages of smaller transformer models for inference workloads and have implications for the sustainable deployment of transformer-based systems at scale.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/southeastcon63549.2026.11476460
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.