MARATTO

article

Enhancing CNN-Based Speech Emotion Recognition with Data Augmentation: A Reproducibility Study

Abstract

A new area of emotional computing called voice Emotion Recognition (SER) aims to extract human emotions from voice data. This study looks into classifying emotions using Convolutional Neural Networks (CNNs) based on voice audio data. To increase dataset diversity, we applied several data augmentation techniques, including pitch shifting, time-stretching, noise addition, and time-shifting. We used Mel-Frequency Cepstral Coefficients (MFCCs) and wave plots to capture important features. A CNN model was trained on the processed data, showing promising accuracy in telling apart different emotional states. We assess the model’s usefulness with confusion matrices and performance metrics, though we still face challenges in distinguishing between acoustically similar emotions. This paper demonstrates how deep learning techniques can help create emotion-aware systems that enhance human-computer interaction.

Research topics

  • Emotion and Mood Recognition
  • Music and Audio Processing
  • Speech Recognition and Synthesis

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icoa66896.2025.11236879

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.