Loading...
Search for:
hadoop
0.071 seconds
Fault Tolerance in Cloud Storage Systems Using Erasure Codes
, M.Sc. Thesis Sharif University of Technology ; Miremadi, Ghassem (Supervisor)
Abstract
International Data Company (IDC) has reported, at the end of 2020, the total amount of digital data stored in the entire world will reach 40 thousand Exabytes. The idea of accessing this volume of data, anywhere at any time by exploiting commodity hardware, led into the introduction of cloud storage. The abounded rate and variety of failures in the equipment used in cloud storage systems, placed fault tolerance, at top of the challenges in these systems. HDFS layer in Hadoop has provided cloud with reliable storage. Replication is the conventional method to protect data against failures in HDFS. But the storage overhead is a big deal and therefore designers are tending towards erasure codes....
Investigating Performance Bottlenecks for Efficient Implementation of MapReduce in Hadoop
, M.Sc. Thesis Sharif University of Technology ; Goudarzi, Maziar (Supervisor)
Abstract
Ever-increasing development and growth of information volume is an unprecedented phenomenon. Analyzing and saving such enormous volume of information calls for innovative ideas capable of processing and managing this information. One of the successful projects in this regard done by Apache is known as Hadoop.Hadoop is a popular open-source implementationof MapReduce processing schemefor analysis of large datasets. The heart of Hadoop is MapReduce that is a parallel programming model for data processing on clusters. To handle storage resources across the cluster, Hadoop employs a distributeduser-level filesystem. The Hadoop Distributed File System (HDFS) is written in Java and is designed for...
Network-aware Key Partitioner for Efficient MapReduce Computation
, M.Sc. Thesis Sharif University of Technology ; Goudarzi, Maziar (Supervisor)
Abstract
MapReduce and its open source implementation, Hadoop, are the prevailing platforms for big data processing. MapReduce is a simple programming model for performing large computational problems in large-scale distributed systems. This model consists of two major phases: Map and Reduce. Between these two main phases, partitioner part is embedded which distributes produced keys by Map tasks among Reduce tasks When the amount of keys and their associated values, which are called intermediate data, is huge, this part has significant impact on execution time of Reduce tasks, and consequently, completion time of jobs. In this paper, we present a network and resource aware key partitioner to decrease...