MARATTO

article

Comparative Analysis of Energy Consumption and Carbon Footprint in Automatic Speech Recognition Systems: A Case Study Comparing Whisper and Google Speech-to-Text

Abstract

This study investigates the energy consumption and carbon footprint of two prominent automatic speech recognition (ASR) systems: OpenAI’s Whisper and Google’s Speech-to-Text API. We evaluate both local and cloud-based speech recognition approaches using a public Kaggle dataset of 20,000 short audio clips in Urdu, utilizing CodeCarbon, PyJoule, and PowerAPI for comprehensive energy profiling. As a result of our analysis, we expose some substantial differences between the two systems in terms of energy efficiency and carbon emissions, with the cloud-based solution showing substantially lower environmental impact despite comparable accuracy. We discuss the implications of these findings for sustainable AI deployment and minimizing the ecological footprint of speech recognition technologies.

Research topics

  • Speech Recognition and Synthesis
  • Music and Audio Processing
  • Modeling, Simulation, and Optimization

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.3390/cmsf2025010006

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.