SPEECH EMOTION RECOGNITION USING A HYBRID CNN–SVM FRAMEWORK WITH THREE-CHANNEL MEL–DELTA–DELTA² SPECTROGRAM FEATURESID: 3829 Abstract :Speech Emotion Recognition (SER) Has Become A Key Area Of Study In Affective Computing.It Allo Ws Smart Systems To Understand Human Emotions By Analyzin G Spoken Words, Which Is Useful In How People Interact With Computers., Healthcare, Virtual Assistants, And Smart Communication Systems. Accurate Recognition Of Emotions Remains Challenging Due To Variations In Speech Characteristics, Speaker Dependency, And Acoustic Conditions. This Paper Presents A Hybrid SER Framework That Combines Deep Transfer Learning With Machine Learning To Improve Emotion Classification Performance. Speech Recordings Are First Preprocessed Through Mono Conversion, Resampling, And Amplitude Normalization. Each Audio Sample Is Converted Into A Threechannel Representation That Includes Mel Spectrogram, Delta, And Delta-delta Features.. This Representation Preserves Both Spectral And Temporal Information. A Pre-trained ResNet18 Network Is Fine-tuned To Extract Meaningful Deep Features, Which Are Then Classified Using A Multi-class Support Vector Machine (SVM) With A Radial Basis Function Kernel. The Proposed Framework Is Tested On The RAVDESS Emotional Speech Dataset, Which Includes Six Different Emotion Categories. The Experimental Results Show That The Overall Classification Accuracy Is 93.04%, And The Precision, Recall, And F1-score Are 0.94, 0.93, And 0.93 Respectively. A MATLAB-based Graphical User Interface (GUI) Is Also Developed To Provide An End-to-end Platform For Audio Preprocessing, Spectrogram Visualization, And Real-time Emotion Prediction. Overall, The Proposed Hybrid CNN–SVM Framework Demonstrates Reliable Emotion Recognition While Maintaining Computational Efficiency, Making It Suitable For Practical Affective Computing Applications. |
Published:13-8-2026 Issue:Vol. 26 No. 8 (2026) Page Nos:857-864 Section:Articles License:This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. How to CiteMeghana Kuruba, Dr. P. Ramana Reddy, SPEECH EMOTION RECOGNITION USING A HYBRID CNN–SVM FRAMEWORK WITH THREE-CHANNEL MEL–DELTA–DELTA² SPECTROGRAM FEATURES , 2026, International Journal of Engineering Sciences and Advanced Technology, 26(8), Page 857-864, ISSN No: 2250-3676. |