Speech Emotion Recognition
A 3D CNN application that recognizes seven emotions from speech, achieving 86.49% accuracy on a 777-sample test set.

Overview
Speech Emotion Recognition classifies seven emotional states from speech: angry, disgust, fearful, happy, neutral, sad, and surprised. The project combines a Deep Learning training pipeline with a Streamlit interface where users can upload an audio file or record their voice and receive a prediction.
Problem
Emotion in speech varies with the speaker, language, recording conditions, and background noise. Class imbalance also makes a model more likely to favor some emotions and less able to generalize between classes with similar acoustic characteristics.
Role
I served as the team lead for the project.
Solution
- Combined 5,147 audio samples from RAVDESS, TESS, SAVEE, and EmoDB into seven consistent emotion labels.
- Converted each audio clip into a normalized mel-spectrogram and a 3D tensor so the model could learn spatial and temporal features.
- Applied audio augmentation to simulate noise and varied recording conditions.
- Trained a 3D CNN and evaluated it across stratified training, validation, and test sets.
- Built a Streamlit interface for inference from uploaded files or live voice recordings.
Technical decisions
The model uses a 3D CNN instead of treating each spectrogram as an independent image, allowing it to learn frequency patterns and their changes over time. This architecture offered enough capacity for complex audio patterns while remaining practical within the team's training constraints.
The data was split 70/15/15 with label stratification. Class-weighted Focal Loss replaced standard cross-entropy to reduce the effect of class imbalance and focus training on harder emotions. Early stopping, model checkpointing, and learning-rate reduction on plateaus were used to control overfitting and retain the model with the best validation accuracy.
Results
The model achieved 86.49% accuracy on a 777-sample test set, correctly classifying 667 samples. Among the seven emotions, surprised achieved the highest F1-score at 0.93. The evaluation also identified two areas for improvement: recall for happy and precision for disgust.
Lessons learned
This was my first time building a Deep Learning model and the starting point of my first project. The process showed me how much more there is to learn and explore through future experiments.
Keep exploring