International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1705959

1705959PublishedVol 7 · Issue 12

Comparative Analysis of Persistence Storage Levels in Spark with Case Study

Thet Hsu Aung Aye Myat Myat Paing

Subject area: Science,Engineering and Technology  ·  Area of research: High Performance Computing

Abstract

This study conducts a comparative analysis of the training times for Long Short-Term Memory (LSTM) networks on Apache Spark, evaluating three different persistence storage levels: Disk-Only, Memory-Disk, and Memory-Only. The analysis is performed with and without a proposed sampling algorithm designed to address the issue of imbalanced datasets. The study specifically focuses on the Credit Card Fraud Detection dataset across varying dataset sizes. The results indicate that the Memory-Only storage level achieves the shortest training times. Considering this case study, the amount of the dataset influences the storage level selection. Therefore, this dataset indicates that memory_only is the best. Memory_disk_only is the second-best option, if memory becomes insufficient due to the growing dataset. Therefore, the choice of storage level affects performance, especially in memory usage and computation speed. Furthermore, the application of the sampling algorithm significantly enhances model performance metrics, including precision, recall, and F1-score, particularly in scenarios involving imbalanced data. These findings provide crucial insights for improving LSTM training on large-scale imbalanced datasets, highlighting the importance of selecting appropriate storage configurations and preprocessing techniques in big data environments.

Keywords

Long Short-Term Memory, Disk-Only, Memory-Only, Memory-Disk

How to cite this paper

Thet Hsu Aung, Aye Myat Myat Paing "Comparative Analysis of Persistence Storage Levels in Spark with Case Study" Iconic Research And Engineering Journals Volume 7 Issue 12 2024 Page 304-313
Thet Hsu Aung, Aye Myat Myat Paing "Comparative Analysis of Persistence Storage Levels in Spark with Case Study" Iconic Research And Engineering Journals, vol. 7, no. 12, Jun. 2024
Thet Hsu Aung, Aye Myat Myat Paing (2024). Comparative Analysis of Persistence Storage Levels in Spark with Case Study. Iconic Research And Engineering Journals, 7(12).
Thet Hsu Aung, Aye Myat Myat Paing "Comparative Analysis of Persistence Storage Levels in Spark with Case Study" Iconic Research And Engineering Journals, vol. 7, no. 12, Jun. 2024.
@article{1705959,
      author = {Thet Hsu Aung, Aye Myat Myat Paing},
      title = {Comparative Analysis of Persistence Storage Levels in Spark with Case Study},
      journal = {Iconic Research And Engineering Journals},
      year = {2024},
      volume = {7},
      number = {12},
      pages = {304-313},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1705959.pdf},
      abstract = {This study conducts a comparative analysis of the training times for Long Short-Term Memory (LSTM) networks on Apache Spark, evaluating three different persistence storage levels: Disk-Only, Memory-Disk, and Memory-Only. The analysis is performed with and without a proposed sampling algorithm designed to address the issue of imbalanced datasets. The study specifically focuses on the Credit Card Fraud Detection dataset across varying dataset sizes. The results indicate that the Memory-Only storage level achieves the shortest training times. Considering this case study, the amount of the dataset influences the storage level selection. Therefore, this dataset indicates that memory_only is the best. Memory_disk_only is the second-best option, if memory becomes insufficient due to the growing dataset. Therefore, the choice of storage level affects performance, especially in memory usage and computation speed. Furthermore, the application of the sampling algorithm significantly enhances model performance metrics, including precision, recall, and F1-score, particularly in scenarios involving imbalanced data. These findings provide crucial insights for improving LSTM training on large-scale imbalanced datasets, highlighting the importance of selecting appropriate storage configurations and preprocessing techniques in big data environments.},
      keywords = {Long Short-Term Memory, Disk-Only, Memory-Only, Memory-Disk},
      month = {June},
  }