Saturday, 20 April 2013

Map Reduce: A really simple introduction

Ever since google published its research paper on map reduce, you have been hearing about it. Here and there. If you have uptil now considered map-reduce a mysterious buzzword, and ignored it, Know that its not. The basic concept is really very simple. and in this tutorial I try to explain it in the simplest way that I can. Note that I have intentionally missed out some deeper details to make it really friendly to a beginner.

Chapter 1: Your CEO’s Strange itch:

Imagine this. You work in a really big company. Your company is planning to launch the next big “Blogging platform”. Tommorow morning you go to your office and there’s a mail from your CEO regarding a new work:
 
Dear  <Your Name>,
 As you know we are building the blogging platform blogger2.com, I need some statistics. I need to find out, Acorss all blogs ever wrriten on blogger.com, how many times 1 character words occur(like 'a', 'I'), How many times two character words occur (like 'be', 'is').. and so on till how many times do ten character words occur.

 I know its a really big job. So, I will assign, all 50,000 employees working in our company to work with you on this for a week.  I am going on a vacation for a week, and its really important that I've this when I return. Good luck.

regds,
The CEO

P.s : and one more thing. Everything has to be done manually, except going to the blog and copy pasting it on notepad. I read somewhere that if you write programs, google can find out about it

Picture yourself in that position for a moment. You have 50,000 people to work for you for a week. And you need to find out the number of 1 character words, No. of 2 character words etc., covering the maximum number of blogs in blogspot. Finally you need to give a report to your CEO with something like this:
  • Occurance of one character words – Around 937688399933
  • Occurance of two chracter words – Around 23388383830753434
  • .. hence forth till 10
If homicide, suicide or resigining the job is not an option, how would you solve it? How would you avoid the chaos of so many people working. How will you co-ordinate those many since the output of one has to be merged with another?
You decide to take leave for the day, go home, sleep over it, and the next day wake up with the greatest Idea ever. “S**t! i wasted a day!”

Chapter 2: Your proclamation: Let there be caste

The next day, You stand with a mike on the dias before 50,000 and proclaim. For a week, you will all be divided into many groups:
  • The Mappers (tens of Thousands of people will be in this group)
  • The Grouper (Assume just one guy for now)
  • The Reducers( Around 10 of em.) and..
  • The Master(That’s you).
Then you talk to each one of the groups.

Chapter 3: Your talk with The Mappers

Each mapper will get a set of 50 blog urls and really Big sheet of paper. Each one of you need to go to each of that url. and for each word in those blogs, write one line on the paper. The format of that line should be the number of characters in the word, then a commna, and then the actual word.

For example, if you find the word “a”, you write “1,a”, in a new line in your paper. since the word “a” has only 1 character. If you find the word “hello”, you write “5,hello” on the new line.

Each take 4 days. So, After 4 days, your sheet might look like this
  • “1,a”
  • “5,hello”
  • “2,if”
  • .. and a million more lines
At the end of the 4th day. each one of you will give your sheet completely filled to the Grouper

Chapter 4: Your talk with the Grouper

I will give you 10 papers. The first paper will be marked 1, the second paper will be marked 2, and so on, till 10.
You collect the output from mappers and for each line in the mapper’s sheet, if it says “1,”, your write the on sheet 1, if it says “2, ”, you write it on sheet two.

For example, if the first line of a mapper’s sheet says “1,a”, you write “a” on sheet 1. if it says “2,if”, your write “if” on sheet 2. If it says “5,hello”, you write hello on sheet 5.

So at the end of your work, the 10 sheets you have might look like this
  • Sheet 1: a, a ,a , I, I , i, a, i, i, i…. millions more
  • Sheet 2: if, of, it, of, of, if, at, im, is,is, of, of … millions more
  • Sheet 3 :the, the, and, for, met, bet, the, the, and, … millions more
  • ..
  • Sheet 10: ……
once you are done, you distribute, each sheet to one reducer. For example sheet 1 goes to reducer 1, sheet 2 goes to reducer 2 and so on.

Chapter 5: Your talk with the Reducers:

Each one of you gets one sheet from the grouper. for each sheet you count the number of words written on it and write it in big bold letters on the back side of the paper.

For ex, if you are reducer 2. You get sheet 2 from the grouper that looks like this:
“Sheet 2: if, of, it, of, of, if, at, im, is,is, of, of …”

You count the number of words on that sheet, say the number of words is 28838380044, You write it on the back side of the paper , in big bold letters and give it to me(the master).

Chapter 6: The controlled Chaos and the climax:

At the end of this process you have 10 sheets, Sheet 1, having the count of the number of words with 1 character on the back side. Sheet2, having the count of the number words with 2 characetrs on the back side. You did it. Genius.
You essentially did map reduce. The greatest advantage in your approach was this
  • The mappers can work independently
  • The reducers can work independently
  • The Grouper can work really fast, because, he din’t have to do any counting of words, all the had to do was to look at the first number and put that word in the appropriate sheet.
The process can be easily applied to other kinds of problems. In such a case :
  • The work of the Master(dividing the work) and the Grouper(Grouping the values by key[the value before commna]), remains the same. This is what any map-reduce library provides.
  • The work of the mappers and reducers differ according to the problem. This is what you should write.
You can optimize this a little bit. And I am skipping those. For example, you don’t even have to mention the words, every where, you could have just written down “x”, instead of the actual word, since in the end, we are just counting. And everything need not happen in a sequence like First Mapper, the Grouper and then reducer. Moreover, one person can be sometime do the job of a mapper and some other time the job of a grouper. Give all this a thought and you will get more answers.
So You solve the biggest challenge ever posed to you.

After a week You collect the sheet of papers from the reducers. The back side of sheet 1 will have the number of occurences of words with 1 character. The back side of sheet 2 will have the number of occurances of words with 2 characters and so on..

You put this information in a excel, Take a printout in a neat sheet of paper and take it to your CEO with a big smile. “Good job “, he says, “put it on the desk, I will take a look at it in a month” :)

Wednesday, 17 April 2013

Confused About Map/Reduce ?



 I was working on some Hadoop stuff recently, and as a total beginner, I found that the Map/Reduce concept was not easy to understand, despite the huge number of tutorials.
The Wordcount example is the ‘Hello World’ of Hadoop, but when I prepared a small presentation for my team, I realized it was not clear enough to explain Map/Reduce in 5 minutes.
As you may already know, the Map/Reduce pattern is a pattern that is very good for embarrassingly parallel algorithms.
Okayyyy but… What is an embarrassingly parallel algorithm?
Answer: It is an algorithm that is very well fit to be executed multiple times in parallel.
Ok then… what is very well suited for a parallel execution?
Answer: Any algorithm that’s working on data that can be isolated.
When writing an application, if you execute multiple occurrences of it at the same time, and they need to access some common data, there will be some clash, and you will have to handles cases like when one occurrence is changing some data while another other is reading it. You’re doingconcurrency.
But if your occurrence is working on some data that no other occurrence will need, then you’re doingparallelism. Obviously you can scale further, since you do not have concurrency issues.
So let’s take an example, let’s say you have a list of cities, and each one has two attributes : the state it belongs to, and its yearly average temperature. E.g. : San Francisco : {CA, 58}
Now you want to calculate the yearly average temperature BY STATE.
Since you can group cities by state, and calculate the average temperature of a state without caring about cities of other states, you have a great embarrassingly parallel algorithm candidate.
If you wanted to do it sequentially, you would start with an empty list of yearly state average temperatures. Then you would iterate through the list of cities, and for each city, look at the state, then update the relevant yearly state average temperature.
Fortunately, it’s very easy to do it in parallel instead.
Let’s have a look at this map:
This is a map of India. There are several states : MP, CG, OR… And several cities, each one having {State, City average temperature} as value.
We want here to calculate the yearly average per state. In order to do that, we should group the city average temperatures by state, then calculate the average of each group.
We don’t really care about the city names, so we will discard those and keep only the state names and cities Temperatures.
Now we have only the data we need, and we can regroup the temperatures values by state. We’re going to get a list of temperatures averages for each state.
At this point, we have the data in good shape to actually do the maths… All we have to do is to calculate the average temperature for each state
That wasn’t hard.
We had some input data. We did a little regrouping, then we did the calculation. And all this could be executed in parallel (One parallel task for each state).
Well… That was Map/Reduce!
Let’s do it again
Map/Reduce has 3 stages : Map/Shuffle/Reduce
The Shuffle part is done automatically by Hadoop, you just need to implement the Map and Reduce parts.
You get input data as <Key,Value>  for the Map part.
In this example, the Key is the City name, and the Value is the set of  attributes : State and City yearly average temperature.
Since you want to regroup your temperatures by state, you’re going to get rid of the city name, and the State will become the Key, while the Temperature will become the Value.
Now, the shuffle task will run on the output of the Map task. It is going to group all the values by Key, and you’ll get a List<Value>
And this is what the Reduce task will get as input : the Key, List<Value> from the Shuffle task.
The Reduce task is the one that does the logic on the data, in our case this is the calculation of the State yearly average temperature.
And that’s what we will get as final output
This is how the data is shaped across Map/Reduce:
Mapper <K1, V1> —> <K2, V2>
Reducer <K2, List<V2>> —><K3, V3>
I hope this helped makes things a bit clearer about Map/Reduce, if you’re interested in explanations about Map Reduce v2/YARN, just leave a comment and I’ll post another entry.
PS: The java code for this example can be found here:

Monday, 25 March 2013

Use the Windows Key for the “Start” Menu in Ubuntu Linux


Linux distributions like Ubuntu open the main menu with Alt+F1 instead of the Windows key that most new Linux       users would be expecting, but it used to be simple to change the shortcut key. Since Ubuntu 9.10 the process isn’t so obvious, but we’ve got the instructions for you.
Just in case you’re a total newb, here’s the menu we’re talking about:
image

Change the Gnome Main Menu Shortcut Key to the Windows Key

The first thing you’d normally do is head to System –> Preferences –> Keyboard Shortcuts to change out the          shortcut key, but sadly the “Show the panel’s main menu” can’t be assign to the Windows key. You can hit the                     key as much as you want, but it won’t work here.
What you’re going to need to do is either open up a terminal or use the Alt+F2 shortcut key to bring up the Run Application dialog, and then paste in the following:
gconftool-2 --set /apps/metacity/global_keybindings/panel_main_menu --type string "Super_L"
Once you’ve hit the enter key, the Windows key will not only open the main menu, but the Keyboard Shortcuts           panel will be updated with “Super L”, which means the left Windows key.
And there you go.

Friday, 22 March 2013

Uptobox Working Premium Link Generator List 2013



Server 1: Uptobox, Uploading, Crocko, Mediafire, iFile, Cramit
(Registration Required)

Server 2: Uptobox, Crocko, Bayfile, Extabit, Jumbofiles, Netload, Rapidshare, Filefactory, Megashares, Sockshare, Uploaded and more...

Server 3: Uptobox, Crocko, Rapidgator



I will update the list more often once i get more servers, so please come back for the update.


mageia - open-ssh instalation



As many of you face this problem that if install Mageia from DVD
it wouldn't install openssh properly.
So if type ssh and enter in terminal/konsole it will say "bash: ssh: command not found".
For this You have to do
su
urpmi ssh ssh-server

After this
service sshd start
chkconfig --list sshd
chkconfig sshd on

This will enable you to access you computer from other computer though ssh by ip adress(e.g ssh user@192.168.9.98).

Friday, 1 March 2013

Project ideas for Hadoop

If this is for an undergraduate class, I would suggest something that
allows you to get some work in with basic data structures such as
building an inverted index over a few million documents (maybe Wikipedia
pages?). You will also need to get a general feel for Hadoop.
The University of Washington has some really nice project ideas for
their distributed systems class:


If you wanted to tackle something a little more advanced, then you could
take a look at Pete Skomoroch’s article on finding trends with Hadoop
and Hive:



Things to keep in mind:

1.) Hadoop wont be as simple as writing a single Java app
 
2.) There will be some overhead involved in re-writing algorithms in Map Reduce
 
3.) There will also be some overhead involved in setup and maintenance
of the Hadoop Cluster

Take these three things into account when planning how to manage your
time for the project during the semester, semesters can seem a lot
shorter when you spend too much time on things not related to just
implementing and testing your algorithm.


 Good luck!

Thursday, 28 February 2013

HOW TO START WORKING WITH HADOOP : PART 1

You can find countless posts on the same topic over the internet. And most of them are really good. But quite often, newbies face some issues even after doing everything as specified. I was no exception. In fact, many a times, my friends who are just starting their Hadoop journey, call me up and tell me that they are facing some issues even after doing everything in order. So, I thought of writing down the things which worked for me. I am not going in detail as there are many better post that outline everything pretty well. I'll just show how to configure Hadoop on a single Linux box in pseudo distributed mode.

Prerequisites :

1- Sun(Oracle) java must be installed on the machine.
2- ssh must be installed and keypair must be already generated.

NOTE : Ubuntu comes with its own java compiler (i.e OpenJDK), but Sun(Oracle) java is the preferable choice for Hadoop. You can visit this link if you need some help on how to install it.

NOTE : You can visit this link if you want to see how to setup and configure ssh on your Ubuntu box.

Versions used :

1- Linux (Ubuntu 12.04)
2- Java (Oracle java-7)
3- Hadoop (Apache hadoop-1.0.3)
4- OpenSSH_5.9p1 Debian-5ubuntu1, OpenSSL 1.0.1 14

If you have everything in place, start following the steps shown below to configure Hadoop on your machine :

1- Download the stable release of Hadoop (hadoop-1.0.3 at the time of this writing) from the repository and copy it to some convenient location. Say your home directory.

2- Now, right click the compressed file which you have downloaded just now and choose extract here. This will create the hadoop-1.0.3 folder inside your home directory. We'll call this location as  HADOOP_HOME hereafter. So, your HADOOP_HOME=/home/your_username/hadoop-1.0.3

3- Edit the /HADOOP_HOME/conf/hadoop-env.sh file to set the JAVA_HOME variable to point to appropriate jvm.

    export JAVA_HOME=/usr/lib/jvm/java-7-oracle

NOTE : Before moving further, create a directory, hdfs for instance, with sub directories viz. name, data and tmp. We'll use these directories as the values of properties in the configuration files.

NOTE : Change the permissions of the directories created in the previous step to 755. Too open or too close permissions may result in abnormal behavior. Use the following command to do that :

cluster@ubuntu:~$ sudo chmod -R 755 /home/cluster/hdfs/


4- Now, we'll start with the actual configuration process. Hadoop is configured using a set of configuration files present inside the HADOOP_HOME/conf directory. These are xml files having a set of properties in form of key-value pairs. We'll modify the following 3 files for our setup :

    I- HADOOP_HOME/conf/core-site.xml : Add the following lines between the <configuration></configuration> tag -

    <property>
            <name>fs.default.name</name>
            <value>hdfs://localhost:9000</value>
     </property>
     <property>
             <name>hadoop.tmp.dir</name>
             <value>/home/your_username/hdfs/tmp</value>
     </property>

     fs.default.name : This is the URI (protocol specifier, hostname, and port) that describes the NameNode for  the cluster. Each node in the system on which Hadoop is expected to operate needs to know the address of the NameNode.

    hadoop.tmp.dir : A base for temporary directories. Value of this property defaults to the /tmp directory. So, it is always better to set this property to some other location to prevent irregularities.

     II- HADOOP_HOME/conf/hdfs-site.xml : Add the following lines between the <configuration></configuration> tag -

     <property>
             <name>dfs.name.dir</name>
             <value>/home/your_username/hdfs/name</value>
      </property>
      <property>
             <name>dfs.data.dir</name>                  
             <value>/home/your_username/hdfs/data</value>
      </property>
      <property>
             <name>dfs.replication</name>
             <value>1</value>
      </property>

      dfs.name.dir : This is the path on the local file system of the NameNode instance where the NameNode metadata is stored. Defaults to the /tmp directory, if not specified explicitly.

    dfs.data.dir : This is the path on the local file system in which the DataNode instance should store its data. It also defaults to the /tmp directory, if not specified explicitly.

  II- HADOOP_HOME/conf/mapred-site.xml : Add the following lines between the <configuration></configuration> tag -

     <property>
              <name>mapred.job.tracker</name>
              <value>localhost:9001</value>
      </property>

      mapred.job.tracker : host and port at which JobTracker will run.

NOTE : Although there are many properties that can be used and play an important role while working with a large, fully distributed cluster, above shown properties are sufficient enough to set up a pseudo distributed Hadoop cluster on a single machine.

5- The configuration part is over now. And in order to proceed further, we have to format our Hdfs first (like any other file system). Use the following command to do that :

    cluster@ubuntu:~/hadoop-1.0.3$ bin/hadoop namenode -format

If everything was ok, you'll see something like this on your terminal :

12/07/23 05:43:22 INFO namenode.NameNode: STARTUP_MSG: 
/************************************************************
STARTUP_MSG: Starting NameNode
STARTUP_MSG:   host = ubuntu/127.0.0.1
STARTUP_MSG:   args = [-format]
STARTUP_MSG:   version = 1.0.3
STARTUP_MSG:   build = https://svn.apache.org/repos/asf/hadoop/common/branches/branch-1.0 -r 1335192; compiled by 'hortonfo' on Tue May  8 20:31:25 UTC 2012
************************************************************/
12/07/23 05:43:22 INFO util.GSet: VM type       = 64-bit
12/07/23 05:43:22 INFO util.GSet: 2% max memory = 17.77875 MB
12/07/23 05:43:22 INFO util.GSet: capacity      = 2^21 = 2097152 entries
12/07/23 05:43:22 INFO util.GSet: recommended=2097152, actual=2097152
12/07/23 05:43:22 INFO namenode.FSNamesystem: fsOwner=cluster
12/07/23 05:43:22 INFO namenode.FSNamesystem: supergroup=supergroup
12/07/23 05:43:22 INFO namenode.FSNamesystem: isPermissionEnabled=true
12/07/23 05:43:22 INFO namenode.FSNamesystem: dfs.block.invalidate.limit=100
12/07/23 05:43:22 INFO namenode.FSNamesystem: isAccessTokenEnabled=false accessKeyUpdateInterval=0 min(s), accessTokenLifetime=0 min(s)
12/07/23 05:43:22 INFO namenode.NameNode: Caching file names occuring more than 10 times 
12/07/23 05:43:22 INFO common.Storage: Image file of size 113 saved in 0 seconds.
12/07/23 05:43:23 INFO common.Storage: Storage directory /home/cluster/hdfs/name has been successfully formatted.
12/07/23 05:43:23 INFO namenode.NameNode: SHUTDOWN_MSG: 
/************************************************************
SHUTDOWN_MSG: Shutting down NameNode at ubuntu/127.0.0.1
************************************************************/
cluster@ubuntu:~/hadoop-1.0.3$

6- Once the formatting is done, start the NameNode, Secondary NameNode and DataNode daemons using the command shown below :   

     cluster@ubuntu:~/hadoop-1.0.3$ bin/start-dfs.sh


This will emit the following lines on the terminal :


starting namenode, logging to /home/cluster/hdfs/logs/hadoop-cluster-namenode-ubuntu.out

localhost: starting datanode, logging to home/cluster/hdfs/logs/hadoop-cluster-datanode-ubuntu.outlocalhost: starting secondarynamenode, logging to /home/cluster/hdfs/logs/hadoop-cluster-secondarynamenode-ubuntu.out
cluster@ubuntu:~/hadoop-1.0.3$

7- To start the JobTracker and Tasktracker daemons use :

      cluster@ubuntu:~/hadoop-1.0.3$ bin/start-mapred.sh

This will emit the following lines on the terminal :

starting jobtracker, logging to /home/cluster/hdfs/logs/hadoop-cluster-jobtracker-ubuntu.out
localhost: starting tasktracker, logging to
/home/cluster/hdfs/logs/hadoop-cluster-tasktracker-ubuntu.out
cluster@ubuntu:~/hadoop-1.0.3$

NOTE : To check if everything is working fine or not, we'll use JPS command (OpenJDK must be installed for this) :
cluster@ubuntu:~/hadoop-1.0.3$ jps
12537 Jps
12042 SecondaryNameNode
12173 JobTracker
11783 DataNode
11487 NameNode
12421 TaskTracker

NOTE : Hadoop also provides a web interface using which we can monitor our cluster. Point your web browser to http://localhost:50070 to see the NameNode status and to http://localhost:50030 to see the MapReduce status.



You can find all the information about your Hdfs from this page. You can even browse the file system and download files from here.