# MapReduce Explained!



*This article is a part of  [Big-Data](https://hashnode.com/series/big-data-ckjvopi6v0904bds13cobg7ts)  series in which I'll be posting stuff related to Big data tech stack. Most of the articles will be short and to the point.*

If you're following along in the Big-Data series where we talked about  [Hadoop](https://shivamahirao.hashnode.dev/hadoop-101) and one of the most crucial part of Hadoop aka The **heart** of Hadoop in the previous blog post -  [*HDFS Explained*](https://shivamahirao.hashnode.dev/hdfs-explained).  Now it's time to decipher the brain of Hadoop .i.e.. MapReduce.

## MapReduce

MapReduce is a programming model and an associated implementation for processing and generating big data sets with a parallel, distributed algorithm on a cluster. A MapReduce program is composed of a map procedure, which performs filtering and sorting, and a reduce method, which performs a summary operation.  
( [*Wikipedia*](https://en.wikipedia.org/wiki/MapReduce) )

### They say:

![mapreduce-is-not-easy.jpg](https://cdn.hashnode.com/res/hashnode/image/upload/v1610822702059/ez4xBCJPe.jpeg)


But let's try to understand it the easy way!

- MapReduce is the computing layer of Hadoop.
- It is responsible for processing large amount of data parallelly/ distributedly.
- The computation is done parallelly on the physical servers .i.e. commodity hardware.
- The MapReduce is designed to work on large set of data.

> 
Since the Hadoop stores data distributedly via HDFS so it allows MapReduce to do the processing parallelly over the cluster.
 
## How MapReduce works?

- It divides the work into independent tasks, lots of nodes work in parallel this way high processing power is achieved.
- MapReduce works on the concept of Data locality:

> **Data Locality**:
Instead of moving data close to the code, Code (computation) should be moved close to the data.

## MR job and Task

### MR Job
- Client is responsible to create/code a MR(Map Reduce) job.
- Client creates the job, configures it and submits it to Master.
- Master sends copy of this MR job(code) to each node .i.e. slaves.

> 
**MapReduce supports various languages** : C, C++, Java, Ruby, Python etc.

![1_XBwye59MGy8l9TJ-NgOipg.jpeg](https://cdn.hashnode.com/res/hashnode/image/upload/v1610822553448/dhJtctXSz.jpeg)


### Task
- An execution of a mapper and Reducer on a slice of data.
- Master divides the work into tasks and schedules it on the slave.

## MapReduce Daemons:

![slide-17-1024.jpg](https://cdn.hashnode.com/res/hashnode/image/upload/v1610822069678/XKtug6XkN.jpeg)


## Mapper and Reducer


![neeson-map-reduce.jpg](https://cdn.hashnode.com/res/hashnode/image/upload/v1610822595676/_mX-ihDpH.jpeg)

### Mapper : 
- A program that performs Map phase using Map() function in MapReduce job.
- Takes data in the form of Key value pairs.
- Ideally 10-100 maps/node.

### Reducer :
- A program that performs Reduce phase using Reduce() function in MapReduce job.
- Takes data in the form of Key value pairs from mapper.
- Accepts o/p of mapper as the input.

> MapReduce data elements are always structured as key value(K,V) pairs.


## Phases of MapReduce

![fadf32ab-b857-4d22-a334-c989b5bafdea.png](https://cdn.hashnode.com/res/hashnode/image/upload/v1610822200724/YTyhXEoLo.png)

Let's understand the concept covered in image in 3 steps:

1. **Map stage** divides the input data into key value pairs.
2. **Shuffle/Sort** is a stage which makes sure that all the data is stored in same node.
3. **Reduce stage** returns single value for each key.

Well this is enough for one go! 

So now finally you can say!!

![download.jfif](https://cdn.hashnode.com/res/hashnode/image/upload/v1610822435806/PPC5laFZG.jpeg)


%%[twitterfollowbutton]
