Home / Current Issue / Paper 1701535
Hadoop MapReduce Performance Improvement In Distributed System
Subject area: Science,Engineering and Technology · Area of research: Computer Engineering
Abstract
MapReduce is currently a parallel computing framework for distributed processing of large-scale data intensive application. The most important performance metric is job execution time but it can be seriously impacted by straggler machines. Speculative execution is a common approach for this problem by backing up slow tasks on alternative machines. Some schedulers with speculative execution have been proposed but they have some weaknesses: (i) they cannot calculate the progress rate accurately because the progress scores of the phases are set to constant values which may be totally different for heterogeneous environment, (ii) they define the stragglers by specifying a static threshold value which calculates the temporal difference between an individual task and the average task progression. To get the better performance, this paper proposes an algorithm identifying the stragglers by the more accurate progress of each job based on its own historical information and using a dynamic threshold value adjusting the continuously varying environment automatically
References
[1] H. Li, Introduction to Big Data, New York, October 31, 2015.
[2] M. Zaharia, A. Konwinski, A. D. Joseph, R. Katz, I.Stoica, “Improving MapReduce Performance in Heterogeneous Environments”, 8th USENIX Symposium on Operating Systems Design and Implementation, March 2009, pp. 29-42.
[3] J. Dean and S. Ghemawat "MapReduce: simplified data processing on large clusters", Communications of the ACM, vol.51, January 2008, pp.107-113.
[4] Q. Chen, C. Liu, and Z. Xiao, “Improving MapReduce Performance Using Smart Speculative Execution Strategy”, IEEE Transactions on Computers, Volume 63, Issue 4, April 2014, pp. 1-14.
[5] P. Garraghan, X. Ouyang, R. Yang, D. McKee, J. Xu, “Straggler Root-Cause and Impact Analysis for Massive-scale Virtualized Cloud Data centres”, IEEE Transactions on Services Computing, 2016, pp. 1-13.
[6] Speculative Execution. [Online]. Available: http://hadoopinrealworld.com/speculative- execution/
[7] “Apache hadoop, http://hadoop.apache.org/”.
[8] M. Isard, M. Budiu, Y. Yu, A. Birrell, and D. Fetterly, “Dryad: distributed data-parallel programs from sequential building blocks,” in Proc. of the 2nd ACM SIGOPS/EuroSys European Conference on Computer Systems 2007, ser. EuroSys ’07, 2007.
[9] Q. Chen, M. Guo, Q. Deng, L. Zheng, S. Guo, Y. Shen, “SAMR: A Self -adaptive MapReduce Scheduling Algorithm In Heterogeneous Environment”, 10th IEEE International Conference on Computer and Information Technology, 2010, pp. 2736-2743.
[10] Xiaoyu Sun, “An Enhanced Self-adaptive MapReduce Scheduling Algorithm”, Master Thesis, University of Nebraska, Lincoln, 2012.
[11] Q. Chen, M. Guo, Q. Deng, L. Zheng, S. Guo, Y. Shen, “HAT: history-based auto-tuning MapReduce in heterogeneous environments”, The Journal of Supercomputing, June 2013, Volume 64, Issue 3, pp 1038–1054.
[12] X. Ouyang, P. Garraghan, D. Mckee, P. Townend, J. Xu, “Straggler Detection in Parallel Computing Systems through Dynamic Threshold Calculation”, 30th IEEE International Conference on Advanced Information Networking and Applications”, 2016, pp.414-421.
How to cite this paper
@article{1701535,
author = {Saw Mya Nandar, Thaint Zarli Myint},
title = {Hadoop MapReduce Performance Improvement In Distributed System},
journal = {Iconic Research And Engineering Journals},
year = {2019},
volume = {3},
number = {2},
pages = {487-493},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1701535.pdf},
abstract = {MapReduce is currently a parallel computing framework for distributed processing of large-scale data intensive application. The most important performance metric is job execution time but it can be seriously impacted by straggler machines. Speculative execution is a common approach for this problem by backing up slow tasks on alternative machines. Some schedulers with speculative execution have been proposed but they have some weaknesses: (i) they cannot calculate the progress rate accurately because the progress scores of the phases are set to constant values which may be totally different for heterogeneous environment, (ii) they define the stragglers by specifying a static threshold value which calculates the temporal difference between an individual task and the average task progression. To get the better performance, this paper proposes an algorithm identifying the stragglers by the more accurate progress of each job based on its own historical information and using a dynamic threshold value adjusting the continuously varying environment automatically},
month = {August},
}