Home / Current Issue / Paper 1705959
Comparative Analysis of Persistence Storage Levels in Spark with Case Study
Subject area: Science,Engineering and Technology · Area of research: High Performance Computing
Abstract
This study conducts a comparative analysis of the training times for Long Short-Term Memory (LSTM) networks on Apache Spark, evaluating three different persistence storage levels: Disk-Only, Memory-Disk, and Memory-Only. The analysis is performed with and without a proposed sampling algorithm designed to address the issue of imbalanced datasets. The study specifically focuses on the Credit Card Fraud Detection dataset across varying dataset sizes. The results indicate that the Memory-Only storage level achieves the shortest training times. Considering this case study, the amount of the dataset influences the storage level selection. Therefore, this dataset indicates that memory_only is the best. Memory_disk_only is the second-best option, if memory becomes insufficient due to the growing dataset. Therefore, the choice of storage level affects performance, especially in memory usage and computation speed. Furthermore, the application of the sampling algorithm significantly enhances model performance metrics, including precision, recall, and F1-score, particularly in scenarios involving imbalanced data. These findings provide crucial insights for improving LSTM training on large-scale imbalanced datasets, highlighting the importance of selecting appropriate storage configurations and preprocessing techniques in big data environments.
Keywords
Long Short-Term Memory, Disk-Only, Memory-Only, Memory-Disk
References
[1] J. Kim, D. Lee, and C. Yoo, “Big data analytics on in-memory database for smart manufacturing”, Journal of Manufacturing Systems, 2019.
[2] R. Y. Park, J. H. Kim, J. Y. Yoo, “A hybrid machine learning approach to financial early warning systems”, Expert Systems with Applications, 2020.
[3] X. Huang, J. Liu, W. Wu, and J. Xie, “Efficient distributed deep learning using edge computing for IoT”, IEEE Internet of Things Journal, 2017.
[4] D. Jiang, Z. Xu, Z. Lin, and J. Guo, “A parallel data mining algorithm on Hadoop for big data. Journal of Parallel and Distributed Computing, 2018.
[5] M. Zaharia, M. Chowdhury, J. Franklin, S. Shenker, and I. Stoica, “Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing”, USENIX NSDI, 2012.
[6] T. M. Khoshgoftaar, J. Van Hulse, and A. Napolitano, “Comparing boosting and bagging techniques with noisy and imbalanced data”, IEEE Transactions on Systems, Man, and Cybernetics, 2013.
[7] https://www.geeksforgeeks.org/stratified-random-sampling-an-overview/#what-is-stratified-random-sampling
[8] https://statusneo.com/solving-data-skewness-in-apache-spark-techniques-and-best-practices/
[9] https://sparktpoint.com/spark-persistence-storage-levels/
[10] https://sparkbyexamples.com/spark/spark-difference-between-cache-and-persist/
[11] https://medium.com/data-engineer/cache-vs-persist-in-apache-spark-a-detailed-comparison-cce43c529599
[12] https://www.kaggle.com/datasets/mlgulb/creditcardfraud
[13] https://medium.com/apache-spark-performance-with-caching-and-persistence
[14] ABFGKLMN[\^qrstvxy€”•–º»ñåñåñåñÕȽ³§š³§‘ƒuuƒii]Q]EhÖ98h^~ 6�OJQJhÖ98hÖ986�OJQJhÖ98hYzõ6�OJQJhÖ98h£}Ë6�OJQJhÖ98hÖ986�H*OJQJhÖ98h^~ 6�H*OJQJh^~ H*OJQJh^~ hÖ98OJPJQJh^~ hÖ98H*OJQJhÖ98OJPJQJh^~ h^~ OJQJh^~ h^~ CJOJQJ*h°ãh°ãCJ(OJPJQJhñ`µCJ(OJPJQJh°ãh°ãCJ(OJPJQJMNs»¼½]^kl<=òåл©©ŸŸŸŠzŸŸŸ$ &Fdë¤a$gd˜g$„ñÿ„‹dì¤:]„ñÿ^„‹a$gd«u $¤a$gdÖ98„ñÿ„ dð¤]„ñÿ^„ gd^~ $„ñÿ„ dð¤]„ñÿ^„ a$gdÖ98$„`„öÿdë¤^„``„öÿa$gd˜g$dð¤a$gd^~ $dݤa$gd˜g»¼½ÅÇ‹ ½ è é HX «ï]½R [Some characters in this reference could not be displayed correctly — please refer to the published PDF for the full reference.]
[15] !ôåÓ¿«——————ƒ—o««[««G¿G¿Ó&hÖ98hÖ985�6�CJOJPJQJaJ&hÖ98hñ`µ5�6�CJOJPJQJaJ&hÖ98hIpL5�6�CJOJPJQJaJ&hÖ98hƒH5�6�CJOJPJQJaJ&hÖ98h.¯5�6�CJOJPJQJaJ&hÖ98h˜‹5�6�CJOJPJQJaJ&hÖ98h^~ 5�6�CJOJPJQJaJ"hÖ98h^~ 5�6�CJOJQJaJh^~ h˜gCJOJQJaJh^~ CJOJQJaJ!\]^jkll;<=T]^=hÖ98h4)CJOJQJaJ&hÖ98hñ`µ5�6�CJOJPJQJaJhÖ98hñ`µCJOJQJaJhÖ98héS§CJOJQJaJhÖ98CJOJQJaJh˜gh˜gCJOJQJaJh˜gh^~ CJOJQJaJ"h˜gh£L¯5�6�CJOJQJaJh^~ 5�6�CJOJQJaJ"hÖ98h«u5�6�CJOJQJaJ$=\]45XYרlm<=]^pq0$1$á'â',,Ð.õõõõõõõõõõõõõõõõÞõõõõõõõõhÖ98hÞKçCJOJQJaJhÖ98hŸ ‡CJOJQJaJhÖ98h*oCJOJQJaJhÖ98h˜TúCJOJQJaJhÖ98hòZCJOJQJaJhÖ98h{wŒCJOJQJaJhÖ98hðU�CJOJQJaJhÖ98hà/¢CJOJQJaJhÖ98CJOJQJaJhÖ98hÞ9CJOJQJaJhÖ98héS§CJOJQJaJhÖ98hée™CJOJQJaJ¶ÌÑÒ\]^opqŠŽ•–q!Ó!Ô!ê!ì!q#/$0$1$R$V$]$_$Ò$Ó$ý%þ%&&1&–&—&Æ&Ó&ñâñÓǸ§˜Œ}}}}}}}}}}}˜Œnnnnnnnnnnnnnn_hÖ98hðU�CJOJQJaJhÖ98h®AgCJOJQJaJhÖ98h”M¡CJOJQJaJhÖ98CJOJQJaJhÖ98h^~ CJOJQJaJ hÖ98húJÂCJOJPJQJaJhÖ98hÖ98CJOJQJaJhà/¢CJOJQJaJhÖ98h*oCJOJQJaJhÖ98hNK[CJOJQJaJhÖ98hBãCJOJQJaJ%Ó&à'á'â'î'ò'û'ü'4)5)â)e*f*â+ [Some characters in this reference could not be displayed correctly — please refer to the published PDF for the full reference.]
How to cite this paper
@article{1705959,
author = {Thet Hsu Aung, Aye Myat Myat Paing},
title = {Comparative Analysis of Persistence Storage Levels in Spark with Case Study},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {7},
number = {12},
pages = {304-313},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1705959.pdf},
abstract = {This study conducts a comparative analysis of the training times for Long Short-Term Memory (LSTM) networks on Apache Spark, evaluating three different persistence storage levels: Disk-Only, Memory-Disk, and Memory-Only. The analysis is performed with and without a proposed sampling algorithm designed to address the issue of imbalanced datasets. The study specifically focuses on the Credit Card Fraud Detection dataset across varying dataset sizes. The results indicate that the Memory-Only storage level achieves the shortest training times. Considering this case study, the amount of the dataset influences the storage level selection. Therefore, this dataset indicates that memory_only is the best. Memory_disk_only is the second-best option, if memory becomes insufficient due to the growing dataset. Therefore, the choice of storage level affects performance, especially in memory usage and computation speed. Furthermore, the application of the sampling algorithm significantly enhances model performance metrics, including precision, recall, and F1-score, particularly in scenarios involving imbalanced data. These findings provide crucial insights for improving LSTM training on large-scale imbalanced datasets, highlighting the importance of selecting appropriate storage configurations and preprocessing techniques in big data environments.},
keywords = {Long Short-Term Memory, Disk-Only, Memory-Only, Memory-Disk},
month = {June},
}