International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1708497

1708497 Vol 6 · Issue 6 Download Paper

Impacts and Outcomes of Using Dropout Layers to Mitigate Overfitting in Convolutional Neural Networks

Rajat Gupta Rakesh Jindal Amisha Naik K L Ganatre

Subject area: Science,Engineering and Technology  ·  Area of research: Neural Networks

Abstract

Convolutional Neural Networks (CNNs) have demonstrated high accuracy in various computer vision tasks such as image classification, object detection, and facial recognition. However, these models are prone to overfitting?especially when they are highly complex and trained on limited data. Overfitting hampers a model?s ability to generalize to unseen data, making regularization essential in deep learning. One of the most effective and commonly used regularization techniques is dropout, which involves randomly deactivating a subset of neurons during each training iteration. This process reduces the risk of neurons becoming overly reliant on specific training features, thereby promoting robustness and better generalization. In this study, we empirically examine the impact of dropout layers within CNN architectures. Our focus is on understanding how different dropout rates influence training behavior, generalization capabilities, and overall model performance. We conduct experiments using well-known image classification datasets under a range of dropout configurations. Across all trials, our findings consistently show that incorporating dropout leads to lower overfitting, improved validation accuracy, and enhanced performance on unseen data. These results underscore the importance of integrating dropout into CNN designs, particularly when working with smaller datasets. Our analysis also reveals the critical balance required when selecting a dropout rate, as both excessively high and low rates can impair model effectiveness through underfitting or insufficient regularization. Ultimately, our study affirms dropout as a key technique for improving the robustness and reliability of deep learning models in computer vision.

Keywords

Convolutional Neural Networks, dropout, overfitting, image classification, regularization, generalization.

References

[1] [1] J. M. Ahn, J. Kim, and K. Kim, “Ensemble Machine Learning of Gradient Boosting (XGBoost, LightGBM, CatBoost) and Attention-Based CNN-LSTM for Harmful Algal Blooms Forecasting,” Toxins (Basel), vol. 15, no. 10, pp. 1–15, Oct. 2023, doi: 10.3390/toxins15100608.

[2] [2] A. Anton, N. F. Nissa, A. Janiati, N. Cahya, and P. Astuti, “Application of Deep Learning Using Convolutional Neural Network (CNN) Method For Women’s Skin Classification,” Scientific Journal of Informatics, vol. 8, no. 1, pp. 144–153, May 2021, doi: 10.15294/sji.v8i1.26888.

[3] [3] M. Ghislieri, G. L. Cerone, M. Knaflitz, and V. Agostini, “Long short-term memory (LSTM) recurrent neural network for muscle activity detection,” J Neuroeng Rehabil, vol. 18, no. 1, pp. 1–15, Dec. 2021, doi: 10.1186/s12984-021-00945-w.

[4] [4] N. Mohd, H. Singhdev, and D. Upadhyay, “Text Classfication Using CNN and CNN-LSTM,” Webology, vol. 18, no. 4, pp. 2440–2446, 2021, doi: 10.29121/web/v18i4/149.

[5] [5] S. Saadah, K. M. Auditama, A. A. Fattahila, F. I. Amorokhman, A. Aditsania, and A. A. Rohmawati, “Implementation of BERT, IndoBERT, and CNN-LSTM in Classifying Public Opinion about COVID-19 Vaccine in Indonesia,” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 6, no. 4, pp. 648–655, Aug. 2022, doi: 10.29207/resti.v6i4.4215.

[6] [6] Y. Gong, “STL: A Signed and Truncated Logarithm Activation Function for Neural Networks,” arxiv.org, vol. 14, no. 8, pp. 1–5, Jul. 2021, doi: 10.48550/arXiv.2307.16389.

[7] [7] H. Kaur, “Sentiment Analysis Of User Review Text Through Cnn And Lstm Methods,” PalArch’s Journal of Archaeology of Egypt / Egyptology, vol. 17, no. 12, pp. 290–306, 2020.

[8] [8] P. N. Anggreyani and W. Maharani, “Hoax Detection Tweets of the COVID-19 on Twitter Using LSTM-CNN with Word2Vec,” Jurnal Media Informatika Budidarma, vol. 6, no. 4, pp. 2432–2437, Oct. 2022, doi: 10.30865/mib.v6i4.4564.

[9] [9] L. Khan, A. Amjad, K. M. Afaq, and H. T. Chang, “Deep Sentiment Analysis Using CNN-LSTM Architecture of English and Roman Urdu Text Shared in Social Media,” Applied Sciences (Switzerland), vol. 12, no. 5, pp. 1–18, Mar. 2022, doi: 10.3390/app12052694.

[10] [10] A. N. Ulfah, M. K. Anam, N. Y. S. Munti, S. Yaakub, and M. B. Firdaus, “Sentiment Analysis of the Convict Assimilation Program on Handling Covid-19,” JUITA: Jurnal Informatika, vol. 10, no. 2, pp. 209–216, 2022, doi: 10.30595/juita.v10i2.12308.

[11] [11] S. Sarica and J. Luo, “Stopwords in technical language processing,” PLoS One, vol. 16, no. 8, pp. 1–13, Aug. 2021, doi: 10.1371/journal.pone.0254937.

[12] [12] H. Alshalabi, S. Tiun, N. Omar, F. N. AL-Aswadi, and K. Ali Alezabi, “Arabic light-based stemmer using new rules,” Journal of King Saud University - Computer and Information Sciences, vol. 34, no. 9, pp. 6635–6642, Oct. 2022, doi: 10.1016/j.jksuci.2021.08.017.

[13] [13] K. Maharana, S. Mondal, and B. Nemade, “A review: Data pre-processing and data augmentation techniques,” Global Transitions Proceedings, vol. 3, no. 1, pp. 91–99, Jun. 2022, doi: 10.1016/j.gltp.2022.04.020.

[14] [14] J. T. Hancock and T. M. Khoshgoftaar, “Survey on categorical data for neural networks,” J Big Data, vol. 7, no. 1, pp. 1–41, Dec. 2020, doi: 10.1186/s40537-020-00305-w.

[15] [15] L. Jen and Y.-H. Lin, “A Brief Overview of the Accuracy of Classification Algorithms for Data Prediction in Machine Learning Applications,” Journal of Applied Data Sciences, vol. 2, no. 3, pp. 84–92, 2021, doi: 10.47738/jads.v2i3.38.

[16] [16] S. A. Hicks et al., “On evaluation metrics for medical applications of artificial intelligence,” Sci Rep, vol. 12, no. 1, pp. 1–9, Dec. 2022, doi: 10.1038/s41598-022-09954-8.

[17] [17] S. Orozco-Arias, J. S. Piña, R. Tabares-Soto, L. F. Castillo-Ossa, R. Guyot, and G. Isaza, “Measuring performance metrics of machine learning algorithms for detecting and classifying transposable elements,” Processes, vol. 8, no. 6, pp. 1–18, Jun. 2020, doi: 10.3390/PR8060638.

[18] [18] Esfahani, Shirin Nasr, and Shahram Latifi. “A Survey of State-of-The-Art GAN-Based Approaches to Image Synthesis.” 9th International Conference on Computer Science, Engineering and Applications (CCSEA 2019), 13 July 2019, csitcp.com/paper/9/99csit06.pdf, https://doi.org/10.5121/csit.2019.90906.

[19] [19] Nabati, R., & Qi, H. (2019). "RRPN: Radar Region Proposal Network for Object Detection in Autonomous Vehicles." 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan, 2019, pp. 3093-3097, doi: 10.1109/ICIP.2019.8803392.

[20] [20] Rawat, W., & Wang, Z. (2017). "Deep Convolutional Neural Networks for Image Classification: A Comprehensive Review." Neural Computation, 29(9), pp. 2352-2449, Sept. 2017, doi: 10.1162/neco_a_00990.

[21] [21] Wang, Weibin, et al. “Medical Image Classification Using Deep Learning.” Intelligent Systems Reference Library, 19 Nov. 2019, pp. 33–51, https://doi.org/10.1007/978-3-030-32606-7_3.

[22] [22] Alom, Md Zahangir, et al. “The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches.” ArXiv:1803.01164 [Cs], 12 Sept. 2018, arxiv.org/abs/1803.01164.

[23] [23] Frid-Adar, Maayan, et al. “GAN-Based Synthetic Medical Image Augmentation for Increased CNN Performance in Liver Lesion Classification.” Neurocomputing, vol. 321, Dec. 2018, pp. 321–331, https://doi.org/10.1016/j.neucom.2018.09.013.

[24] [24] Karp, Rafal, and Zaneta Swiderska-Chadaj. Automatic Generation of Graphical Game Assets Using GAN. 13 July 2021, https://doi.org/10.1145/3477911.3477913.

[25] [25] L. Jiao and J. Zhao, "A Survey on the New Generation of Deep Learning in Image Processing," in IEEE Access, vol. 7, pp. 172231-172263, 2019, doi: 10.1109/ACCESS.2019.2956508.

[26] [26] L. Wang, W. Chen, W. Yang, F. Bi and F. R. Yu, "A State-of-the-Art Review on Image Synthesis With Generative Adversarial Networks," in IEEE Access, vol. 8, pp. 63514-63537, 2020, doi: 10.1109/ACCESS.2020.2982224.

[27] [27] Shorten, Connor, and Taghi M. Khoshgoftaar. “A Survey on Image Data Augmentation for Deep Learning.” Journal of Big Data, vol. 6, no. 1, 6 July 2019, journalofbigdata.springeropen.com/articles/10.1186/s40537-019-0197-0, https://doi.org/10.1186/s40537-019-0197-0.

[28] [28] Kayalibay, Baris, et al. “CNN-Based Segmentation of Medical Imaging Data.” ArXiv:1701.03056 [Cs], 25 July 2017, arxiv.org/abs/1701.03056.

[29] [29] Jain, M., & Shah, A. (2022). Machine Learning with Convolutional Neural Networks (CNNs) in Seismology for Earthquake Prediction. Iconic Research and Engineering Journals, 5(8), 389–398. https://www.irejournals.com/paper-details/1707057

[30] [30] Kaushik, P., & Jain, M. A Low Power SRAM Cell for High Speed Applications Using 90nm Technology. Csjournals. Com, 10. https://www.csjournals.com/IJEE/PDF10-2/66.%20Puneet.pdf

[31] [31] Kaushik, P., & Jain, M. (2018). Design of low power CMOS low pass filter for biomedical application. International Journal of Electrical Engineering & Technology (IJEET), 9(5).

[32] [32] Kumar, Y., Saini, S., & Payal, R. (2020). Comparative Analysis for Fraud Detection Using Logistic Regression, Random Forest and Support Vector Machine. SSRN Electronic Journal.

[33] [33] Höppner, S., Baesens, B., Verbeke, W., & Verdonck, T. (2020). Instance-Dependent Cost-Sensitive Learning for Detecting Transfer Fraud. arXiv preprint arXiv:2005.02488.

[34] [34] Niu, X., Wang, L., & Yang, X. (2019). A Comparison Study of Credit Card Fraud Detection: Supervised versus Unsupervised. arXiv preprint arXiv:1904.10604.

[35] [35] Bhat, N. (2019). Fraud detection: Feature selection-over sampling. Kaggle. Retrieved from https://www.kaggle.com/code/nareshbhat/fraud-detection-feature-selection-over-sampling

[36] [36] InsiderFinance Wire. (2021). Logistic regression: A simple powerhouse in fraud detection. Medium. Retrieved from https://wire.insiderfinance.io/logistic-regression-a-simple-powerhouse-in-fraud-detection-15ab984b2102

[37] [37] Olaitan, V. O. (2020). Feature-based selection technique for credit card fraud detection. Master’s Thesis, National College of Ireland. Retrieved from https://norma.ncirl.ie/5122/1/olaitanvictoriaolanlokun.pdf

[38] [38] Raymaekers, J., Verbeke, W., & Verdonck, T. (2021). Weight-of-evidence 2.0 with shrinkage and spline-binning. arXiv preprint arXiv:2101.01494. Retrieved from https://arxiv.org/abs/2101.01494

[39] [38] Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2017). Credit card fraud detection: A realistic modeling and a novel learning strategy. IEEE Transactions on Neural Networks and Learning Systems, 29(8), 3784–3797. https://doi.org/10.1109/TNNLS.2017.2736643

[40] [40] Carcillo, F., Dal Pozzolo, A., Le Borgne, Y. A., Caelen, O., Mazzer, Y., & Bontempi, G. (2019). Scarff: A scalable framework for streaming credit card fraud detection with spark. Information Fusion, 41, 182–194. https://doi.org/10.1016/j.inffus.2017.09.005

[41] [41] West, J., & Bhattacharya, M. (2016). Intelligent financial fraud detection: A comprehensive review. Computers & Security, 57, 47–66. https://doi.org/10.1016/j.cose.2015.09.005

[42] [42] Zareapoor, M., & Shamsolmoali, P. (2015). Application of credit card fraud detection: Based on bagging ensemble classifier. Procedia Computer Science, 48, 679–685. https://doi.org/10.1016/j.procs.2015.04.201

[43] [43] Patel, H., & Zaveri, M. (2011). Credit card fraud detection using neural network. International Journal of Innovative Research in Computer and Communication Engineering, 1(2), 1–6. https://www.ijircce.com/upload/2011/october/1_Credit.pdf

[44] [44] Puneet Kaushik, Mohit Jain , Gayatri Patidar, Paradayil Rhea Eapen, Chandra Prabha Sharma (2018). Smart Floor Cleaning Robot Using Android. International Journal of Electronics Engineering. https://www.csjournals.com/IJEE/PDF10-2/64.%20Puneet.pdf

[45] [45] Duman, E., & Ozcelik, M. H. (2011). Detecting credit card fraud by genetic algorithm and scatter search. Expert Systems with Applications, 38(10), 13057–13063. https://doi.org/10.1016/j.eswa.2011.04.102

[46] [46] Puneet Kaushik, Mohit Jain. “A Low Power SRAM Cell for High Speed ApplicationsUsing 90nm Technology.” Csjournals.Com 10, no. 2 (December 2018): 6.https://www.csjournals.com/IJEE/PDF10-2/66.%20Puneet.pdf

[47] [47] Jain, M., & Srihari, A. (2021). Comparison of CAD detection of mammogram with SVM and CNN. IRE Journals, 8(6), 63-75. https://www.irejournals.com/formatedpaper/1706647.pdf

[48] [48] Kaushik, P., Jain, M., & Jain, A. (2018). A pixel-based digital medical images protection using genetic algorithm. International Journal of Electronics and Communication Engineering, 31-37. http://www.irphouse.com/ijece18/ijecev11n1_05.pdf

[49] [49] Kaushik, P., Jain, M., & Shah, A. (2018). A Low Power Low Voltage CMOS Based Operational Transconductance Amplifier for Biomedical Application. https://ijsetr.com/uploads/136245IJSETR17012-283.pdf

[50] [50] Jain, M., & Shah, A. (2022). Machine Learning with Convolutional Neural Networks (CNNs) in Seismology for Earthquake Prediction. Iconic Research and Engineering Journals, 5(8), 389–398. https://www.irejournals.com/paper-details/1707057

[51] [51] Kaushik, P., & Jain, M. (2018). Design of low power CMOS low pass filter for biomedical application. International Journal of Electrical Engineering & Technology (IJEET), 9(5).

How to cite this paper

Rajat Gupta, Rakesh Jindal, Amisha Naik, K L Ganatre "Impacts and Outcomes of Using Dropout Layers to Mitigate Overfitting in Convolutional Neural Networks" Iconic Research And Engineering Journals Volume 6 Issue 6 2022 Page 392-407
Rajat Gupta, Rakesh Jindal, Amisha Naik, K L Ganatre "Impacts and Outcomes of Using Dropout Layers to Mitigate Overfitting in Convolutional Neural Networks" Iconic Research And Engineering Journals, vol. 6, no. 6, Dec. 2022
Rajat Gupta, Rakesh Jindal, Amisha Naik, K L Ganatre (2022). Impacts and Outcomes of Using Dropout Layers to Mitigate Overfitting in Convolutional Neural Networks. Iconic Research And Engineering Journals, 6(6).
Rajat Gupta, Rakesh Jindal, Amisha Naik, K L Ganatre "Impacts and Outcomes of Using Dropout Layers to Mitigate Overfitting in Convolutional Neural Networks" Iconic Research And Engineering Journals, vol. 6, no. 6, Dec. 2022.
@article{1708497,
      author = {Rajat Gupta, Rakesh Jindal, Amisha Naik, K L Ganatre},
      title = {Impacts and Outcomes of Using Dropout Layers to Mitigate Overfitting in Convolutional Neural Networks},
      journal = {Iconic Research And Engineering Journals},
      year = {2022},
      volume = {6},
      number = {6},
      pages = {392-407},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1708497.pdf},
      abstract = {Convolutional Neural Networks (CNNs) have demonstrated high accuracy in various computer vision tasks such as image classification, object detection, and facial recognition. However, these models are prone to overfitting?especially when they are highly complex and trained on limited data. Overfitting hampers a model?s ability to generalize to unseen data, making regularization essential in deep learning. One of the most effective and commonly used regularization techniques is dropout, which involves randomly deactivating a subset of neurons during each training iteration. This process reduces the risk of neurons becoming overly reliant on specific training features, thereby promoting robustness and better generalization. In this study, we empirically examine the impact of dropout layers within CNN architectures. Our focus is on understanding how different dropout rates influence training behavior, generalization capabilities, and overall model performance. We conduct experiments using well-known image classification datasets under a range of dropout configurations. Across all trials, our findings consistently show that incorporating dropout leads to lower overfitting, improved validation accuracy, and enhanced performance on unseen data. These results underscore the importance of integrating dropout into CNN designs, particularly when working with smaller datasets. Our analysis also reveals the critical balance required when selecting a dropout rate, as both excessively high and low rates can impair model effectiveness through underfitting or insufficient regularization. Ultimately, our study affirms dropout as a key technique for improving the robustness and reliability of deep learning models in computer vision.},
      keywords = {Convolutional Neural Networks, dropout, overfitting, image classification, regularization, generalization.},
      month = {December},
  }