Home / Current Issue / Paper 1705305
Optimizing Sentiment Analysis in Hindi Poetry: A Hybrid Model Unifying Deep Learning, Machine Learning, and Metaheuristic Techniques
Subject area: Science,Engineering and Technology · Area of research: Artificial intelligence
Abstract
Sentiment analysis, an automated computational methodology employed for the investigation and assessment of sentiments, emotions, and feelings conveyed in comments, feedback, or critiques, utilizes machine learning techniques to discern text patterns proficiently. This research leverages supervised machine learning, specifically exploring its application in the sentiment analysis of Hindi poetry-based text through the validation of model feasibility and accuracy using the Hindi Poetry Sentiment Corpus. The study delves into the examination of prevalent supervised machine learning techniques, including Multinomial Naive Bayes, Logistic Regression, and Random Forest, alongside deep learning methodologies such as Long Short-Term Memory and Convolutional Neural Networks. To evaluate classifier performance comprehensively, standard datasets are utilized, and metrics such as precision, recall, F1-score, RoC curve, accuracy, running time, and k-fold cross-validation are employed. This analytical approach yields valuable insights into the efficacy of diverse deep learning techniques, aiding practitioners in selecting suitable methods tailored to their specific applications. Furthermore, the investigation incorporates the application of the metaheuristic-based Grey Wolf Optimization technique to discern optimal features from pre-processed data. The genesis of "deep learning" (DL) in artificial neural network research is acknowledged, wherein word vectors trained by Word2Vec are utilized for the input layer (IL) and input into the CNN-LSTM joint model. Subsequently, the output of the joint model undergoes weighting and summation through self-attention before entering the SoftMax classifier, facilitating the emotion classification of the text. Rigorous comparative experiments validate the utility of the proposed model, demonstrating its superior performance over three comparison models [CNN, LSTM, CNN-LSTM] across various evaluation indices. Comparisons with other machine learning techniques, including Random Forest, Logistic Regression, Naive Bayes, CNN, and LSTM, reveal notable accuracies. Specifically, Random Forest, Naive Bayes, CNN, and LSTM achieve accuracies of 87.75%, 85.54%, 91.46%, and 88.72%, respectively. Notably, the proposed ensemble hybrid model attains the highest classification accuracy of 95.54%, precision of 91.44%, recall of 89.63%, and F-score of 90.87%, showcasing its efficacy in sentiment analysis applications.
Keywords
Hindi poetry-based text sentiment analysis, Machine Learning, Deep Learning, Grey Wolf Optimization, natural language processing, CNN-LSTM multi-feature fusion.
References
[1] Vilares, D., Alonso, M. A., & Gómez-Rodríguez, C. (2017). Supervised sentiment analysis in multilingual environments. Information Processing & Management, 53(3), 595-607.
[2] Mujahid, M.; Lee, E.; Rustam, F.; Washington, P.B.; Ullah, S.; Reshi, A.A.; Ashraf, I. Sentiment Analysis and Topic Modeling on Tweets about Online Education during COVID-19. Appl. Sci. 2021, 11, 8438. https://doi.org/10.3390/app11188438.
[3] Balahur, A., & Turchi, M. (2014). Comparative experiments using supervised learning and machine translation for multilingual sentiment analysis. Computer Speech & Language, 28(1), 56-75.
[4] Kim, E. (2006). Reasons and motivations for code-mixing and code-switching. Issues in EFL, 4(1), 43-61.
[5] Singh, Vinay, Aman Varshney, Syed Sarfaraz Akhtar, Deepanshu Vijay, and Manish Shrivastava. (2018)” Aggression detection on social media text using deep neural networks.” Proceedings of the 2nd Workshop on Abusive Language Online (ALW2): 43-50.
[6] Baroi, S. J., Singh, N., Das, R., & Singh, T. D. (2020, December). NITS-Hinglish-SentiMix at SemEval-2020 Task 9: Sentiment Analysis for Code-Mixed social media Text Using an Ensemble Model. In Proceedings of the Fourteenth Workshop on Semantic Evaluation (pp. 1298-1303).
[7] Si, S., Datta, A., Banerjee, S., & Naskar, S. K. (2019, July). Aggression detection on multilingual social media text. In 2019 10th International Conference on Computing, Communication and Networking Technologies (ICCCNT) (pp. 1-5). IEEE.
[8] Wu, Q., Wang, P., & Huang, C. (2020). MeisterMorxrc at SemEval2020 Task 9: Fine-tune bert and multitask learning for sentiment analysis of code-mixed tweets. arXiv preprint arXiv:2101.03028.
[9] Bhange, M., & Kasliwal, N. (2020). HinglishNLP: Fine-tuned Language Models for Hinglish Sentiment Detection. arXiv preprint arXiv:2008.09820.
[10] Parikh, A., Bisht, A. S., & Majumder, P. (2020, December). IRLab_DAIICT at SemEval-2020 Task 9: Machine Learning and Deep Learning Methods for Sentiment Analysis of Code-Mixed Tweets. In Proceedings of the Fourteenth Workshop on Semantic Evaluation (pp. 1265-1269).
[11] Kumar, V., Pasari, S., Patil, V. P., & Seniaray, S. (2020, July). Machine Learning based Language Modelling of Code-Switched Data. In 2020 International Conference on Electronics and Sustainable Communication Systems (ICESC) (pp. 552-557). IEEE.
[12] Dahiya, A., Battan, N., Shrivastava, M., & Sharma, D. M. (2019, August). Curriculum Learning Strategies for Hindi-English Code-Mixed Sentiment Analysis. In International Joint Conference on Artificial Intelligence (pp. 177-189). Springer, Cham.
[13] Singh, P., & Lefever, E. (2020, May). Sentiment Analysis for Hinglish Code-mixed Tweets by means of Cross-lingual Word Embedding‟s. In Proceedings of the The 4th Workshop on Computational Approaches to Code Switching (pp. 45-51).
[14] R. D Endsuy, “Sentiment Analysis between VADER and EDA for the US Presidential Election 2020 on Twitter Datasets”, Journal of Applied Data Sciences 2 (2021) 8.
[15] M. Bibi, W. Aziz, M. Almaraashi, I. H. Khan, M. S. A. Nadeem & N. Habib, “A Cooperative Binary-Clustering Framework Based on Majority Voting for Twitter Sentiment Analysis”, IEEE Access 8 (2020) 68580.
[16] R. Cekik & S. Telceken, “A New Classification Method Based on Rough Sets Theory”, Soft Computing 6 (2018) 1881.
[17] A. Jain & V. Jain, “Sentiment Classification Using Hybrid Feature Selection and Ensemble Classifier” Journal of Intelligent & Fuzzy Systems, 4(2021) 221.
[18] A. P. Rodrigues & N. N. Chiplunkar, “A New Big Data Approach for Topic Classification and Sentiment Analysis of Twitter Data”, Evolutionary Intelligence 2 (2019)11.
[19] S. Rani, N. S. Gill & P. Gulia, “Survey of Tools and Techniques for Sentiment Analysis of Social Networking Data”, International journal of Advanced computer Science and applications 12 (2021) 222.
[20] R. Cekik & A. K. Uysal, “A novel filter feature selection method using rough set for short text data”, Expert Systems with Applications 160 (2020) 113691
[21] Chandra R, Krishna A (2021) Covid-19 sentiment analysis via deep learning during the rise of novel cases. Plos one 16(8):e0255615
[22] Lwin MO, Lu J, Sheldenkar A, Schulz PJ, Shin W, Gupta R, Yang Y (2020) Global sentiments surrounding the covid-19 pandemic on twitter: analysis of twitter trends. JMIR Public Health and Surveillance 6(2):e19447
[23] Prabhakar Kaila D, Prasad DrAV et al (2020) Informational flow on twitter–corona virus outbreak– topic modelling approach. International Journal of Advanced Research in Engineering and Technology (IJARET) 11:3
[24] Nemes L, Kiss A (2021) Social media sentiment analysis based on covid-19. J Inform Telecommun 5(1):1–15
[25] Samuel J, Ali GG, Rahman M, Esawi E, Samuel Y et al (2020) Covid-19 public sentiment insights and machine learning for tweets classification. Information 11(6):314
[26] Mittal N, Agarwal B, Chouhan G, Bania N, Pareek P (2013) Sentiment analysis of hindi reviews based on negation and discourse relation. In: Proceedings of the 11th workshop on Asian language resources, pp 45–50
[27] Gupta V, Jain N, Shubham S, Madan A, Chaudhary A, Xin Q (2021) Toward integrated cnn-based sentiment analysis of tweets for scarce-resource language-hindi. Transactions on Asian and Low-Resource Language Information Processing 20(5):1–23
[28] A. M. G. Almeida, R. Cerri, E. C. Paraiso, R. G. Mantovani, and S. Barbon Junior, “Applying multi-label techniques in emotion identification of short texts,” Neurocomputing, vol. 320, no. 3, pp. 35–46, 2018.
[29] H. Liu, M. Shen, J. Zhu, N. Niu, and L. Zhang, “Deep learning based program generation from requirements text: are we there yet?,” IEEE Transactions on Software Engineering, vol. 48, no. 4, pp. 1268–1289, 2020.
[30] W. Zeng, H. Xu, H. Li, and X. Li, “Research on the methodology of correlation analysis of sci-tech literature based on deep learning technology in the big data,” Journal of Database Management, vol. 29, no. 3, pp. 67–88, 2018.
[31] C. N. Dang, M. N. Moreno-García, and F. Prieta, “Hybrid deep learning models for sentiment analysis,” Complexity, vol. 2021, Article ID 9986920, 16 pages, 2021.
How to cite this paper
@article{1705305,
author = {Vinod kumar, Archismita Ghosh, Kandikattu Sai Rachana, Teetas Bhutiya, Sukanya Wattal},
title = {Optimizing Sentiment Analysis in Hindi Poetry: A Hybrid Model Unifying Deep Learning, Machine Learning, and Metaheuristic Techniques},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {7},
number = {6},
pages = {193-207},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1705305.pdf},
abstract = {Sentiment analysis, an automated computational methodology employed for the investigation and assessment of sentiments, emotions, and feelings conveyed in comments, feedback, or critiques, utilizes machine learning techniques to discern text patterns proficiently. This research leverages supervised machine learning, specifically exploring its application in the sentiment analysis of Hindi poetry-based text through the validation of model feasibility and accuracy using the Hindi Poetry Sentiment Corpus. The study delves into the examination of prevalent supervised machine learning techniques, including Multinomial Naive Bayes, Logistic Regression, and Random Forest, alongside deep learning methodologies such as Long Short-Term Memory and Convolutional Neural Networks. To evaluate classifier performance comprehensively, standard datasets are utilized, and metrics such as precision, recall, F1-score, RoC curve, accuracy, running time, and k-fold cross-validation are employed. This analytical approach yields valuable insights into the efficacy of diverse deep learning techniques, aiding practitioners in selecting suitable methods tailored to their specific applications. Furthermore, the investigation incorporates the application of the metaheuristic-based Grey Wolf Optimization technique to discern optimal features from pre-processed data. The genesis of "deep learning" (DL) in artificial neural network research is acknowledged, wherein word vectors trained by Word2Vec are utilized for the input layer (IL) and input into the CNN-LSTM joint model. Subsequently, the output of the joint model undergoes weighting and summation through self-attention before entering the SoftMax classifier, facilitating the emotion classification of the text. Rigorous comparative experiments validate the utility of the proposed model, demonstrating its superior performance over three comparison models [CNN, LSTM, CNN-LSTM] across various evaluation indices. Comparisons with other machine learning techniques, including Random Forest, Logistic Regression, Naive Bayes, CNN, and LSTM, reveal notable accuracies. Specifically, Random Forest, Naive Bayes, CNN, and LSTM achieve accuracies of 87.75%, 85.54%, 91.46%, and 88.72%, respectively. Notably, the proposed ensemble hybrid model attains the highest classification accuracy of 95.54%, precision of 91.44%, recall of 89.63%, and F-score of 90.87%, showcasing its efficacy in sentiment analysis applications.},
keywords = {Hindi poetry-based text sentiment analysis, Machine Learning, Deep Learning, Grey Wolf Optimization, natural language processing, CNN-LSTM multi-feature fusion.},
month = {December},
}