Team Ai
Apppublic

Saikumar1712/speech-emotion-recognition-customer-behavior-analysis

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes
App README

🎙️ Speech Emotion Recognition — CNN-BiLSTM

A deep-learning app that classifies speech audio into one of five emotions: Angry · Fear · Happy · Neutral · Sad

Model

Architecture: CNN-BiLSTM Hybrid (dual-input)

  • —Sequential branch: 1-D CNN → Bidirectional LSTM, fed a (128, 40) MFCC matrix
  • —Scalar branch: Dense network, fed 7 hand-crafted audio features
  • —Outputs concatenated → softmax over 5 emotions

Dataset: CREMA-D (~6,000 utterances)

Files

FileRequiredPurpose
app.py✅Gradio backend
requirements.txt✅Python dependencies
CNN_bilstm_hybrid_model.h5✅Trained Keras model
seq_scaler.pkl⭐StandardScaler for MFCC branch (improves accuracy)
scaler.pkl⭐StandardScaler for scalar branch (improves accuracy)

If the .pkl scalers are missing, the app falls back to per-sample normalisation. Predictions will still vary correctly but accuracy will be lower.