Home / Current Issue / Paper 1704436
Kannada Speech Emotion Recognition Using Ensembling Techniques
Subject area: Science,Engineering and Technology · Area of research: Speech Processing
Abstract
This study explores the development of a speech emotion recognition system for the Kannada language, using a dataset of audio recordings labeled with six emotion categories: happiness, sadness, anger, fear, and neutral. We used a combination of acoustic features and machine learning algorithms, including Mel-frequency cepstral coefficients (MFCCs), to classify emotions in the audio recordings. Our results show that the proposed system achieves an average accuracy of 75% on the Kannada emotion dataset, outperforming existing baseline models. These findings suggest that Kannada speech emotion recognition can be achieved with high accuracy using a combination of acoustic features and machine learning algorithms like RNN, CNN and DBN, paving the way for further research in this area.
Keywords
Speech Emotion Recognition, Mel-Frequency Cepstral Coefficients, Recurrent Neural Network, Deep Belief Network
References
[1] M. S. Likitha, S. R. R. Gupta, K. Hasitha and A. U. Raju, "Speech based human emotion recognition using MFCC," 2017 International Conference on Wireless Communications, Signal Processing and Networking (WiSPNET),2017,pp.2257-2260,doi: 10.1109/WiSPNET.2017.8300161.
[2] Sonawane, Anagha et al. “Sound based human emotion recognition using MFCC & multiple SVM.” 2017 International Conference on Information, Communication, Instrumentation and Control (ICICIC) (2017): 1-4.
[3] Demircan, Semiye & Kahramanli, Humar. (2014). Feature Extraction from Speech Data for Emotion Recognition. Journal of Advances in Computer Networks. 2. 28-30. 10.7763/JACN.2014.V2.76.
[4] S. Mirsamadi, E. Barsoum and C. Zhang, "Automatic speech emotion recognition using recurrent neural networks with local attention," 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 2227-2231, doi: 10.1109/ICASSP.2017.7952552.
[5] Dongdong Li, Jinlin Liu, Zhuo Yang, Linyu Sun, Zhe Wang,”Speech emotion recognition using recurrent neural networks with directional self-attention, Expert Systems with Applications”, Volume 173, 2021, 114683, ISSN 0957-4174, https://doi.org/10.1016/j.eswa.2021.114683.
[6] L. Kerkeni, Y. Serrestou, M. Mbarki, K. Raoof, M. A. Mahjoub, and C. Cleder, "Automatic Speech Emotion Recognition Using Machine Learning", in Social Media and Machine Learning. London, United Kingdom: IntechOpen, 2019 [Online]. Available: https://www.intechopen.com/chapters/65993 doi: 10.5772/intechopen.84856
[7] P. Shi, "Speech emotion recognition based on deep belief network," 2018 IEEE 15th International Conference on Networking, Sensing and Control (ICNSC), 2018, pp. 1-5, doi: 10.1109/ICNSC.2018.8361376.
[8] B. Chen, Q. Yin and P. Guo, "A Study of Deep Belief Network Based Chinese Speech Emotion Recognition," 2014 Tenth International Conference on Computational Intelligence and Security, 2014, pp. 180-184, doi: 10.1109/CIS.2014.148.
[9] H. Zheng and Y. Yang, "An Improved Speech Emotion Recognition Algorithm Based on Deep Belief Network," 2019 IEEE International Conference on Power, Intelligent Computing and Systems (ICPICS), 2019, pp. 493-497, doi: 10.1109/ICPICS47731.2019.8942482.
[10] V. B. Kobayashi and V. B. Calag, "Detection of affective states from speech signals using ensembles of classifiers," IET Intelligent Signal Processing Conference 2013 (ISP 2013), 2013, pp. 1-9, doi: 10.1049/cp.2013.2067.
[11] N. T. Ira and M. O. Rahman, "An Efficient Speech Emotion Recognition Using Ensemble Method of Supervised Classifiers," 2020 Emerging Technology in Computing, Communication and Electronics (ETCCE), 2020, pp. 1-5, doi: 10.1109/ETCCE51779.2020.9350913.
[12] D. Valles and R. Matin, "An Audio Processing Approach using Ensemble Learning for Speech-Emotion Recognition for Children with ASD," 2021 IEEE World AI IoT Congress (AI IoT), 2021, pp. 0055-0061, doi: 10.1109/AIIoT52608.2021.9454174.
How to cite this paper
@article{1704436,
author = {Smrithi Baliga, Sapna H M, Shreyas N, Yogesh Gowda V, Dr Chandrashekar M Patil; Prof. Audre Arlene},
title = {Kannada Speech Emotion Recognition Using Ensembling Techniques},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {6},
number = {11},
pages = {250-255},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1704436.pdf},
abstract = {This study explores the development of a speech emotion recognition system for the Kannada language, using a dataset of audio recordings labeled with six emotion categories: happiness, sadness, anger, fear, and neutral. We used a combination of acoustic features and machine learning algorithms, including Mel-frequency cepstral coefficients (MFCCs), to classify emotions in the audio recordings. Our results show that the proposed system achieves an average accuracy of 75% on the Kannada emotion dataset, outperforming existing baseline models. These findings suggest that Kannada speech emotion recognition can be achieved with high accuracy using a combination of acoustic features and machine learning algorithms like RNN, CNN and DBN, paving the way for further research in this area.},
keywords = {Speech Emotion Recognition, Mel-Frequency Cepstral Coefficients, Recurrent Neural Network, Deep Belief Network},
month = {May},
}