International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1701538

1701538 Vol 3 · Issue 2 Download Paper

Data Popularity - Aware Replication Strategy For Cloud Storage

May Phyo Thu Khine Moe Nwe Kyar Nyo Aye

Subject area: Science,Engineering and Technology  ·  Area of research: Computer Engineering

Abstract

Replication is one of the important roles in cloud storage to improve data availability, fault tolerance and throughput for users and control storage cost. As data access pattern changes every time, the nature of popular files is unpredictable and unstable. Therefore, data popularity is taken into account as an important factor in replication. Data popularity in replication impacts an efficient storage because it is able to reduce waste storage for unpopular files. Also, data locality is a key issue in storage system and this consequence occurs performance overhead of system. Therefore, this paper introduces a replication strategy for cloud storage. The proposed strategy contains two portions; replica popularity and replica placement. First for replica popularity, popularity is taken into account by analyzing the changes in data access pattern. Second for replica placement, replicas are placed and performed on dedicated assigned nodes in order to enhance data locality. The proposed placement algorithm is able to avoid the overloaded problem of nodes by considering the load of nodes such as disk utilization, adjustable disk bandwidth and CPU utilization. This proposed strategy will be efficient for cloud storage.

Keywords

Popularity, Data Locality, Disk Utilization, Adjustable Disk Bandwidth, CPU Utilization

References

[1] A. Hunger and J. Myint, “Comparative Analysis of Adaptive File Replication Algorithms for Cloud Data Storage”, 2014 International Conference on Future Internet of Things and Cloud, 2014.

[2] B. Gong, B. Veeravalli, D. Feng L. Zeng, and Q. Wei, “CDRM: A Cost-Effective Dynamic Replication Management Scheme for Cloud Storage Cluster”, 2010 IEEE International Conference on Cluster Computing, Sep. 2010, pp. 188–196.

[3] C.L. Abad, Yi Lu, R.H. Campbell, “DARE: Adaptive Data Replication for Efficient Cluster Scheduling”, IEEE International Conference on Cluster Computing (CLUSTER 2011), pp.159-168, 2011.

[4] D. Lee, J. Lee, and J. Chung, “Efficient Data Replication Scheme based on Hadoop Distributed File System”, International Journal of Software Engineering and Its Applications Vol. 9, No. 12 (2015), pp. 177-186,2015.

[5] D.M. Bui, S. Hussain, E.N. Huh, and S. Lee, “Adaptive replication managementin hdfs based on supervised learning,” IEEE Transcations on Knowledage and Data Engineering, vol.28, no.6, 2016.

[6] G. Ananthanarayanan et al., “Scarlett: Coping with skewed content popularity in mapreduce clusters,” in Proc. Conf. Comput. Syst. (EuroSys), 2011, pp. 287–300.

[7] H. Gobioff, S. Ghemawat, and S.-T. Leung, “The Google File System”, Proceedings of 19th ACM Symposium on Operating Systems Principles (SOSP 2003), New York, USA, October, 2003.

[8] H. Hardware, and P. Across, “The Hadoop Distributed File System: Architecture and Design”, 2007, pp. 1–14.

[9] H.-P. Chang, R.-S. Chang, and Y.-T. Wang, “A dynamic weighted data replication strategy in data grids”, 2008 IEEE/ACS International Conference on Computer Systems and Applications, Mar. 2008, pp. 414–421.

[10] M. Zaharia, D. Borthakur, J. Sen Sarma, K. Elmeleegy, S. Shenker, and I. Stoica, “Delay scheduling: A simple technique for achieving locality and fairness in cluster scheduling”, In Proceeding of uropean Conference Computer System (EuroSys), 2010.

[11] https://webscope.sandbox.yahoo.com.

[12] Andrew S. Tanenbaum. Modern Operating Systems. Prentice-Hall, 1992.

How to cite this paper

May Phyo Thu, Khine Moe Nwe, Kyar Nyo Aye "Data Popularity - Aware Replication Strategy For Cloud Storage" Iconic Research And Engineering Journals Volume 3 Issue 2 2019 Page 494-500
May Phyo Thu, Khine Moe Nwe, Kyar Nyo Aye "Data Popularity - Aware Replication Strategy For Cloud Storage" Iconic Research And Engineering Journals, vol. 3, no. 2, Aug. 2019
May Phyo Thu, Khine Moe Nwe, Kyar Nyo Aye (2019). Data Popularity - Aware Replication Strategy For Cloud Storage. Iconic Research And Engineering Journals, 3(2).
May Phyo Thu, Khine Moe Nwe, Kyar Nyo Aye "Data Popularity - Aware Replication Strategy For Cloud Storage" Iconic Research And Engineering Journals, vol. 3, no. 2, Aug. 2019.
@article{1701538,
      author = {May Phyo Thu, Khine Moe Nwe, Kyar Nyo Aye},
      title = {Data Popularity - Aware Replication Strategy For Cloud Storage},
      journal = {Iconic Research And Engineering Journals},
      year = {2019},
      volume = {3},
      number = {2},
      pages = {494-500},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1701538.pdf},
      abstract = {Replication is one of the important roles in cloud storage to improve data availability, fault tolerance and throughput for users and control storage cost. As data access pattern changes every time, the nature of popular files is unpredictable and unstable. Therefore, data popularity is taken into account as an important factor in replication. Data popularity in replication impacts an efficient storage because it is able to reduce waste storage for unpopular files. Also, data locality is a key issue in storage system and this consequence occurs performance overhead of system. Therefore, this paper introduces a replication strategy for cloud storage. The proposed strategy contains two portions; replica popularity and replica placement. First for replica popularity, popularity is taken into account by analyzing the changes in data access pattern. Second for replica placement, replicas are placed and performed on dedicated assigned nodes in order to enhance data locality. The proposed placement algorithm is able to avoid the overloaded problem of nodes by considering the load of nodes such as disk utilization, adjustable disk bandwidth and CPU utilization. This proposed strategy will be efficient for cloud storage.},
      keywords = {Popularity, Data Locality, Disk Utilization, Adjustable Disk Bandwidth, CPU Utilization},
      month = {August},
  }