International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1701959

1701959 Vol 3 · Issue 8 Download Paper

Telephone Voice Speaker Recognition Using Mel Frequency Cepstral Coefficients With Cascaded Feed Forward Neural Network

M. F. Franklin Nissy G. Renisha

Subject area: Science,Engineering and Technology  ·  Area of research: Electronics and Communication Engineering

Abstract

Speaker recognition is the process of identification of the person from the characteristics of his voice. It provides service such as database access services, information services and security control for confidential information areas. However, the accurateness of speaker recognition often drops off quickly because of the low-quality speech and sound. To overcome this problem a new speaker recognition model based on Mel frequency cepstral coefficients (MFCC) are used for feature extraction. Feature extraction means that the speech signal is converted into a series of feature vector coefficients. These features only include the information needed to identify the speaker and discarding all other stuff which carries information like background noise, emotion etc. Features extracted from MFCC are given as the input to the Cascaded Feed Forward Neural Network (CFFNN) which identifies the speech signal of the corresponding speaker. MFCC is an efficient way to extract features from the signal and the Mel scale based feature extraction gives better accuracy in the clean and noisy environment.

Keywords

Speaker recognition, Mel Frequency Cepstral Coefficients (MFCC), Cascaded Feed Forward Neural Network (CFFNN), Feed Forward Neural Network (FFNN)

References

[1] Chougala, M. and Kuntoji, S., 2016, March. Novel text independent speaker recognition using LPC based formants. In 2016 International Conference on Electrical, Electronics, and Optimization Techniques (ICEEOT) (pp. 510-513). IEEE.

[2] Li, P., Hu, F., Li, Y. and Xu, Y., 2014, July. Speaker identification using linear predictive cepstral coefficients and general regression neural network. In Proceedings of the 33rd Chinese Control Conference (pp. 4952-4956). IEEE.

[3] 3.Nair, R. and Salam, N., 2014, December. A reliable speaker verification system based on LPCC and DTW. In 2014 IEEE International Conference on Computational Intelligence and Computing Research (pp. 1-4). IEEE.

[4] 4.Ilyas, M.Z., Samad, S.A., Hussain, A. and Ishak, K.A., 2007. Speaker verification using vector quantization and hidden Markov model. In 2007 5th Student Conference on Research and Development (pp. 1-5). IEEE.

[5] 5. Bansal, P., Imam, S.A. and Bharti, R., 2015, October. Speaker recognition using MFCC, shifted MFCC with vector quantization and fuzzy. In 2015 International Conference on Soft Computing Techniques and Implementations (ICSCTI) (pp. 41-44). IEEE.

[6] 6. Chauhan, N. and Chandra, M., 2017, March. Speaker recognition and verification using artificial neural network. In 2017 International Conference on Wireless Communications, Signal Processing and Networking (WiSPNET) (pp. 1147-1149). IEEE.

[7] Chauhan, N., Isshiki, T. and Li, D., 2019, February. Speaker Recognition Using LPC, MFCC, ZCR Features with ANN and SVM Classifier for Large Input Database. In 2019 IEEE 4th International Conference on Computer and Communication Systems (ICCCS) (pp. 130-133). IEEE.

[8] Weng, Z., Li, L. and Guo, D., 2010, July. Speaker recognition using weighted dynamic MFCC based on GMM. In 2010 International Conference on Anti-Counterfeiting, Security and Identification (pp. 285-288). IEEE.

[9] Sukhwal, A. and Kumar, M., 2015, October. Comparative study between different classifiers based speaker recognition system using MFCC for noisy environment. In 2015 International Conference on Green Computing and Internet of Things (ICGCIoT) (pp. 955-960). IEEE.

[10] Revathy, A., Shanmugapriya, P. and Mohan, V., 2015, March. Performance comparison of speaker and emotion recognition. In 2015 3rd International Conference on Signal Processing, Communication and Networking (ICSCN) (pp. 1-6). IEEE.

How to cite this paper

M. F. Franklin Nissy, G. Renisha "Telephone Voice Speaker Recognition Using Mel Frequency Cepstral Coefficients With Cascaded Feed Forward Neural Network" Iconic Research And Engineering Journals Volume 3 Issue 8 2020 Page 164-170
M. F. Franklin Nissy, G. Renisha "Telephone Voice Speaker Recognition Using Mel Frequency Cepstral Coefficients With Cascaded Feed Forward Neural Network" Iconic Research And Engineering Journals, vol. 3, no. 8, Feb. 2020
M. F. Franklin Nissy, G. Renisha (2020). Telephone Voice Speaker Recognition Using Mel Frequency Cepstral Coefficients With Cascaded Feed Forward Neural Network. Iconic Research And Engineering Journals, 3(8).
M. F. Franklin Nissy, G. Renisha "Telephone Voice Speaker Recognition Using Mel Frequency Cepstral Coefficients With Cascaded Feed Forward Neural Network" Iconic Research And Engineering Journals, vol. 3, no. 8, Feb. 2020.
@article{1701959,
      author = {M. F. Franklin Nissy, G. Renisha},
      title = {Telephone Voice Speaker Recognition Using Mel Frequency Cepstral Coefficients With Cascaded Feed Forward Neural Network},
      journal = {Iconic Research And Engineering Journals},
      year = {2020},
      volume = {3},
      number = {8},
      pages = {164-170},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1701959.pdf},
      abstract = {Speaker recognition is the process of identification of the person from the characteristics of his voice. It provides service such as database access services, information services and security control for confidential information areas. However, the accurateness of speaker recognition often drops off quickly because of the low-quality speech and sound. To overcome this problem a new speaker recognition model based on Mel frequency cepstral coefficients (MFCC) are used for feature extraction. Feature extraction means that the speech signal is converted into a series of feature vector coefficients. These features only include the information needed to identify the speaker and discarding all other stuff which carries information like background noise, emotion etc. Features extracted from MFCC are given as the input to the Cascaded Feed Forward Neural Network (CFFNN) which identifies the speech signal of the corresponding speaker. MFCC is an efficient way to extract features from the signal and the Mel scale based feature extraction gives better accuracy in the clean and noisy environment.},
      keywords = {Speaker recognition, Mel Frequency Cepstral Coefficients (MFCC), Cascaded Feed Forward Neural Network (CFFNN), Feed Forward Neural Network (FFNN)},
      month = {February},
  }