Loading...

Designing a MapReduce performance model in distributed heterogeneous platforms based on benchmarking approach

Gandomi, A ; Sharif University of Technology | 2020

640 Viewed
  1. Type of Document: Article
  2. DOI: 10.1007/s11227-020-03162-9
  3. Publisher: Springer , 2020
  4. Abstract:
  5. MapReduce framework is an effective method for big data parallel processing. Enhancing the performance of MapReduce clusters, along with reducing their job execution time, is a fundamental challenge to this approach. In fact, one is faced with two challenges here: how to maximize the execution overlap between jobs and how to create an optimum job scheduling. Accordingly, one of the most critical challenges to achieving these goals is developing a precise model to estimate the job execution time due to the large number and high volume of the submitted jobs, limited consumable resources, and the need for proper Hadoop configuration. This paper presents a model based on MapReduce phases for predicting the execution time of jobs in a heterogeneous cluster. Moreover, a novel heuristic method is designed, which significantly reduces the makespan of the jobs. In this method, first by providing the job profiling tool, we obtain the execution details of the MapReduce phases through log analysis. Then, using machine learning methods and statistical analysis, we propose a relevant model to predict runtime. Finally, another tool called job submission and monitoring tool is used for calculating makespan. Different experiments were conducted on the benchmarks under identical conditions for all jobs. The results show that the average makespan speedup for the proposed method was higher than an unoptimized case. © 2020, Springer Science+Business Media, LLC, part of Springer Nature
  6. Keywords:
  7. Modeling ; YARN ; Benchmarking ; Data handling ; Heuristic methods ; Job shop scheduling ; Models ; Scheduling ; Yarn ; Hadoop ; Heterogeneous clusters ; Heterogeneous platforms ; Identical conditions ; Machine learning methods ; Makespan ; Map-reduce ; Mapreduce frameworks ; Learning systems ; MapReduce
  8. Source: Journal of Supercomputing ; Volume 76, Issue 9 , 2020 , Pages 7177-7203
  9. URL: https://link.springer.com/article/10.1007/s11227-020-03162-9