International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1704436

1704436 Vol 6 · Issue 11 Download Paper

Kannada Speech Emotion Recognition Using Ensembling Techniques

Smrithi Baliga Sapna H M Shreyas N Yogesh Gowda V Dr Chandrashekar M Patil Prof. Audre Arlene

Subject area: Science,Engineering and Technology  ·  Area of research: Speech Processing

Abstract

This study explores the development of a speech emotion recognition system for the Kannada language, using a dataset of audio recordings labeled with six emotion categories: happiness, sadness, anger, fear, and neutral. We used a combination of acoustic features and machine learning algorithms, including Mel-frequency cepstral coefficients (MFCCs), to classify emotions in the audio recordings. Our results show that the proposed system achieves an average accuracy of 75% on the Kannada emotion dataset, outperforming existing baseline models. These findings suggest that Kannada speech emotion recognition can be achieved with high accuracy using a combination of acoustic features and machine learning algorithms like RNN, CNN and DBN, paving the way for further research in this area.

Keywords

Speech Emotion Recognition, Mel-Frequency Cepstral Coefficients, Recurrent Neural Network, Deep Belief Network

References

[1] M. S. Likitha, S. R. R. Gupta, K. Hasitha and A. U. Raju, "Speech based human emotion recognition using MFCC," 2017 International Conference on Wireless Communications, Signal Processing and Networking (WiSPNET),2017,pp.2257-2260,doi: 10.1109/WiSPNET.2017.8300161.

[2] Sonawane, Anagha et al. “Sound based human emotion recognition using MFCC & multiple SVM.” 2017 International Conference on Information, Communication, Instrumentation and Control (ICICIC) (2017): 1-4.

[3] Demircan, Semiye & Kahramanli, Humar. (2014). Feature Extraction from Speech Data for Emotion Recognition. Journal of Advances in Computer Networks. 2. 28-30. 10.7763/JACN.2014.V2.76.

[4] S. Mirsamadi, E. Barsoum and C. Zhang, "Automatic speech emotion recognition using recurrent neural networks with local attention," 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2017, pp. 2227-2231, doi: 10.1109/ICASSP.2017.7952552.

[5] Dongdong Li, Jinlin Liu, Zhuo Yang, Linyu Sun, Zhe Wang,”Speech emotion recognition using recurrent neural networks with directional self-attention, Expert Systems with Applications”, Volume 173, 2021, 114683, ISSN 0957-4174, https://doi.org/10.1016/j.eswa.2021.114683.

[6] L. Kerkeni, Y. Serrestou, M. Mbarki, K. Raoof, M. A. Mahjoub, and C. Cleder, "Automatic Speech Emotion Recognition Using Machine Learning", in Social Media and Machine Learning. London, United Kingdom: IntechOpen, 2019 [Online]. Available: https://www.intechopen.com/chapters/65993 doi: 10.5772/intechopen.84856

[7] P. Shi, "Speech emotion recognition based on deep belief network," 2018 IEEE 15th International Conference on Networking, Sensing and Control (ICNSC), 2018, pp. 1-5, doi: 10.1109/ICNSC.2018.8361376.

[8] B. Chen, Q. Yin and P. Guo, "A Study of Deep Belief Network Based Chinese Speech Emotion Recognition," 2014 Tenth International Conference on Computational Intelligence and Security, 2014, pp. 180-184, doi: 10.1109/CIS.2014.148.

[9] H. Zheng and Y. Yang, "An Improved Speech Emotion Recognition Algorithm Based on Deep Belief Network," 2019 IEEE International Conference on Power, Intelligent Computing and Systems (ICPICS), 2019, pp. 493-497, doi: 10.1109/ICPICS47731.2019.8942482.

[10] V. B. Kobayashi and V. B. Calag, "Detection of affective states from speech signals using ensembles of classifiers," IET Intelligent Signal Processing Conference 2013 (ISP 2013), 2013, pp. 1-9, doi: 10.1049/cp.2013.2067.

[11] N. T. Ira and M. O. Rahman, "An Efficient Speech Emotion Recognition Using Ensemble Method of Supervised Classifiers," 2020 Emerging Technology in Computing, Communication and Electronics (ETCCE), 2020, pp. 1-5, doi: 10.1109/ETCCE51779.2020.9350913.

[12] D. Valles and R. Matin, "An Audio Processing Approach using Ensemble Learning for Speech-Emotion Recognition for Children with ASD," 2021 IEEE World AI IoT Congress (AI IoT), 2021, pp. 0055-0061, doi: 10.1109/AIIoT52608.2021.9454174.

How to cite this paper

Smrithi Baliga, Sapna H M, Shreyas N, Yogesh Gowda V, Dr Chandrashekar M Patil; Prof. Audre Arlene "Kannada Speech Emotion Recognition Using Ensembling Techniques" Iconic Research And Engineering Journals Volume 6 Issue 11 2023 Page 250-255
Smrithi Baliga, Sapna H M, Shreyas N, Yogesh Gowda V, Dr Chandrashekar M Patil; Prof. Audre Arlene "Kannada Speech Emotion Recognition Using Ensembling Techniques" Iconic Research And Engineering Journals, vol. 6, no. 11, May. 2023
Smrithi Baliga, Sapna H M, Shreyas N, Yogesh Gowda V, Dr Chandrashekar M Patil; Prof. Audre Arlene (2023). Kannada Speech Emotion Recognition Using Ensembling Techniques. Iconic Research And Engineering Journals, 6(11).
Smrithi Baliga, Sapna H M, Shreyas N, Yogesh Gowda V, Dr Chandrashekar M Patil; Prof. Audre Arlene "Kannada Speech Emotion Recognition Using Ensembling Techniques" Iconic Research And Engineering Journals, vol. 6, no. 11, May. 2023.
@article{1704436,
      author = {Smrithi Baliga, Sapna H M, Shreyas N, Yogesh Gowda V, Dr Chandrashekar M Patil; Prof. Audre Arlene},
      title = {Kannada Speech Emotion Recognition Using Ensembling Techniques},
      journal = {Iconic Research And Engineering Journals},
      year = {2023},
      volume = {6},
      number = {11},
      pages = {250-255},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1704436.pdf},
      abstract = {This study explores the development of a speech emotion recognition system for the Kannada language, using a dataset of audio recordings labeled with six emotion categories: happiness, sadness, anger, fear, and neutral. We used a combination of acoustic features and machine learning algorithms, including Mel-frequency cepstral coefficients (MFCCs), to classify emotions in the audio recordings. Our results show that the proposed system achieves an average accuracy of 75% on the Kannada emotion dataset, outperforming existing baseline models. These findings suggest that Kannada speech emotion recognition can be achieved with high accuracy using a combination of acoustic features and machine learning algorithms like RNN, CNN and DBN, paving the way for further research in this area.},
      keywords = {Speech Emotion Recognition, Mel-Frequency Cepstral Coefficients, Recurrent Neural Network, Deep Belief Network},
      month = {May},
  }