On a Dynamic Data Placement Strategy for Heterogeneous Hadoop Clusters
Document Type
Conference Proceeding
Publication Date
11-9-2018
Abstract
Hadoop is one of the most popular distributed systems for big data computing in both industry and science communities. The default data placement strategy of Hadoop Distributed File System (HDFS), which was initially designed for homogenous environments, may suffer from performance degradation when deployed in heterogeneous clusters comprised of data nodes with disparate computing power and disk capacity, hence undermining the performance of MapReduce applications. In this paper, we use a Grey Forecast model to predict data hotness dynamically and determine an appropriate number of data block replicas on the fly. Based on such information, we further propose a dynamic data placement strategy (DDPS) to decide the best location for new replicas according to their hotness. The proposed method is able to dynamically adjust data replicas stored on each node in a heterogeneous Hadoop cluster and reduce the response time of big data applications. Experimental results on a heterogeneous Hadoop cluster show that DDPS together with the prediction model significantly increases application execution efficiency and improve MapReduce performance over the default HDFS configuration.
Identifier
85058473721 (Scopus)
ISBN
[9781538637784]
Publication Title
2018 International Symposium on Networks Computers and Communications Isncc 2018
External Full Text Location
https://doi.org/10.1109/ISNCC.2018.8530970
Grant
61472320
Fund Ref
Northwest University
Recommended Citation
Liu, Yang; Wu, Chase Q.; Wang, Meng; Hou, Aiqin; and Wang, Yongqiang, "On a Dynamic Data Placement Strategy for Heterogeneous Hadoop Clusters" (2018). Faculty Publications. 8260.
https://digitalcommons.njit.edu/fac_pubs/8260
