Saturday, January 11, 2020

BigData PIG Practicals

BigData PIG Practicals
  • LOAD
  • FILTER
  • FOREACH ... GENERATE
  • SPLIT
  • GROUP
  • JOIN
  • DESCRIBE
  • EXPLAIN
  • ILLUSTRATE
  • DUMP
> pig -x local
> pig -x local [script]
> pig -x hadoop [script]

[cloudera@quickstart ~]$ pwd
/home/cloudera
[cloudera@quickstart ~]$ pig -x local
grunt>

/home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/students.txt
grunt> A = load 'Desktop/Basha/Basha2019/PIG_Practicals/students.txt';
grunt> describe A;
Schema for A unknown.
grunt> dump A;
(John,21,2.89)
(Sally,19,2.56)
(Alice,22,3.76)
(Doug,19,1.98)
(Susan,26,3.25)
(John,35,5.00)
(Doug,40,3.50)
(Alice,22,5.25)

grunt> A = load 'Desktop/Basha/Basha2019/PIG_Practicals/students.txt' AS (name:chararray, age:int, gpa:float);
grunt> describe A;
A: {name: chararray,age: int,gpa: float}
grunt> dump A;
(John,21,2.89)
(Sally,19,2.56)
(Alice,22,3.76)
(Doug,19,1.98)
(Susan,26,3.25)
(John,35,5.0)
(Doug,40,3.5)
(Alice,22,5.25)

grunt> R = filter A by (age>=20);
grunt> dump R;
(John,21,2.89)
(Alice,22,3.76)
(Susan,26,3.25)
(John,35,5.0)
(Doug,40,3.5)
(Alice,22,5.25)

grunt> R = filter A by (age>=20) and (gpa>=3.5);
grunt> dump R;
(Alice,22,3.76)
(John,35,5.0)
(Doug,40,3.5)
(Alice,22,5.25)

grunt> illustrate R;
---------------------------------------------------------
| A     | name:chararray    | age:int    | gpa:float    |
---------------------------------------------------------
|       | John              | 35         | 5.0          |
|       | John              | 21         | 2.89         |
|       | Doug              | 19         | 1.98         |
---------------------------------------------------------
---------------------------------------------------------
| R     | name:chararray    | age:int    | gpa:float    |
---------------------------------------------------------
|       | John              | 35         | 5.0          |
---------------------------------------------------------

grunt> F = foreach A generate age,gpa;
grunt> dump F;
(21,2.89)
(19,2.56)
(22,3.76)
(19,1.98)
(26,3.25)
(35,5.0)
(40,3.5)
(22,5.25)

grunt> G = group A by age;
grunt> dump G;
(19,{(Doug,19,1.98),(Sally,19,2.56)})
(21,{(John,21,2.89)})
(22,{(Alice,22,5.25),(Alice,22,3.76)})
(26,{(Susan,26,3.25)})
(35,{(John,35,5.0)})
(40,{(Doug,40,3.5)})
grunt> describe G;
G: {group: int,A: {(name: chararray,age: int,gpa: float)}}

grunt> H = foreach G generate group,A.name;
grunt> dump H;
(19,{(Doug),(Sally)})
(21,{(John)})
(22,{(Alice),(Alice)})
(26,{(Susan)})
(35,{(John)})
(40,{(Doug)})


grunt> store A into 'Desktop/Basha/Basha2019/PIG_Practicals/outputdir';
Input(s):
Successfully read records from: "file:///home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/students.txt"
Output(s):
Successfully stored records in: "file:///home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/outputdir"

grunt> store H into 'Desktop/Basha/Basha2019/PIG_Practicals/outputdir2' using PigStorage('|');
Input(s):
Successfully read records from: "file:///home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/students.txt"
Output(s):
Successfully stored records in: "file:///home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/outputdir2"
19|{(Doug),(Sally)}
21|{(John)}
22|{(Alice),(Alice)}
26|{(Susan)}
35|{(John)}
40|{(Doug)}


[cloudera@quickstart ~]$ pig -x local Desktop/Basha/Basha2019/PIG_Practicals/students.pig
Input(s):
Successfully read records from: "file:///home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/students.txt"
Output(s):
Successfully stored records in: "file:///home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/outputdir3"
19|{(Doug),(Sally)}
21|{(John)}
22|{(Alice),(Alice)}
26|{(Susan)}
35|{(John)}
40|{(Doug)}

Wednesday, January 1, 2020

BigData Pig Concepts

BigData PIG Concepts

- It is a Data flow language.
- PIG allows us to describe how datasets are filtered, combined, split and delivered from source to final destination.
- PIG commands are internally translated into MapReduce jobs.
  • LOAD
  • FILTER
  • FOREACH ... GENERATE
  • SPLIT
  • GROUP
  • JOIN
- PIG will work on any source of tuples. (Not like SQL where schema specified before data load)
- PIG can use complex, nested data structures.
- PIG supports Streaming of data, bulk read + writes. It also describes a series of operations.

- PIG script defines a logical plan for executing a workflow. No actual data is read until execution time.

PIG data structures:


















Relation(DB) Concepts :


















Relation Operations: Relational algebra contains a set of operations that transform one or more relations into other relations.

  • Selection
  • Projection
  • Cartesian Product
  • Extended Projection
  • Aggregation



Query Optimizer: A relational database Query Optimizer considers all the operations needed to produce the result, and finds the most efficient plan to compute it.






















BigData Python Concepts


BigData Python Concepts

- Simple to understand, Easy to learn
- Supports Interactive mode & Script mode
- Interpreted Language
- Case Sensitive, Dynamically Typed
- Object oriented structure
- Platform independent
- Open source
- It has huge number of libraries/packages
- Very useful in Data Science

Important Packages for BigData programming :
  • NumPy
  • pandas
  • SciPy
  • Statsmodels
  • Scikits
  • matplotlib
  • BeautifulSoup

IPython notebook - IPython is a software package in Python that serves as a research notebook. By using IPython, we can write notes as well as perform data analytics in the same file. i.e., write the code and run it from within the notebook itself.

Indentation 
  • Every line is a new line in the Python. There is no end of the line specifier.
  • Indentation = Whitespace at the beginning of the line.
  • Statements which go together must have same indentation. Each such set of statements is called a block.
Variables and Data Structures : 
Build-in data types : Integer, String, Float, Boolean, Date and Time.
Additional data Structures : Lists, Tuples, Dictionary

String Exercises :
course = 'Python for Beginners'print (course)
print (course[0])
print (course[-1])
print (course[2:5])
print (course[2:])
print (course[-10:])
print (course * 2)
print (course + "TEST")

# Different String functionsprint(course.replace("Python", "Jython"))
print(type(course))
print(len(course))
print(course.upper())
print(course.lower())
print(course.count('n'))
print(course.split(' '))
print(course.swapcase())
print(course.strip())
print(course.lstrip())
print(course.rstrip())
print(":".join(course))
print('B' in course)
print(course.find('g'))

empno = input("Enter EMP NO : ")
print(int(empno) + 10)

Lists - Collection of elements of different data types. 
         - List contains items separated by commas and enclosed within square brackets.

Lists Exercises :
emp_list = ['E100','Ravi',1000.00,'Hyderabad','D100']
dept_list = ['D100','Accounts','Delhi']
print(emp_list)
print(dept_list)
print(emp_list[0])
print(emp_list[2:5])
print(emp_list[2:])
print(dept_list * 2)
print(emp_list + dept_list)

print(len(emp_list))
emp_list.append(25)
emp_list.insert(0,'SNO1')
emp_list.extend(['boy1','boy2'])
print(emp_list.index('E100'))
print(emp_list)
emp_list.remove('boy1')
emp_list.pop(1)
print(emp_list)
employees_age_list = [25,20,35,32,21,28,38,45,23,33]
employees_age_list.sort()
print(employees_age_list)
print(len(employees_age_list))
print(max(employees_age_list))
print(min(employees_age_list))
employees_name_list = ['Ravi','Aditya','Giri','Mohanbabu','Madhu']
print(sorted(employees_name_list, key=len))

Dictionary - Consists of Key-Value pairs. enclosed within curly braces.
                   - Key is any Python type, but are usually numbers or strings.
                   - Values can be an arbitrary Python object.

Dictionary Exercises :
# Dictionary Excersizeemp_dictionary={}
print(emp_dictionary)
emp_dictionary['Eno'] = 'E101'emp_dictionary[2] = 'Ename'dept_dictionary = {'dno':101,'dname':'Account','dloc':'Delhi'}
print(emp_dictionary['Eno'])
print(emp_dictionary[2])
print(dept_dictionary)
print(dept_dictionary.keys())
print(dept_dictionary.values())
# empno = input("Enter EMP NO : ")# print(int(empno) + 10)

Conditions - Python supports IF-ELSE, FOR, WHILE....conditions.

IF-Else Exercises:
# IF Condition Excerciseemp_age = int(input('Enter Your Age : '))
if(emp_age >= 25 and emp_age <= 30):
    print('You are eligible for Fresher post')
elif(emp_age < 25):
    print('You are not eligible for the Test')
elif(emp_age > 30 and emp_age < 60):
    print('Your are eligible for Experience post')
else:
    print('Invalid Entry')

emp_name_list = ['Ravi','Giri','John']
emp_name = input('Enter Employee Name : ')
if(emp_name in employees_name_list):
    print(f'{emp_name} present in the employees_name_list')
else:
    print(f'{emp_name} is not present in the emp_name_list')

WHILE Exercises:
# While Condition Excerciseemp_count = 0while(emp_count <= 10):
    print(emp_count)
    emp_count += 1print('End of While loop')

FOR Exercises:
# For Condition Excercisefor dept_no in ('d100','d101','d102'):
    print(dept_no)

for emp_age in range (20,30,2):
    print(emp_age)

emp_age = [22, 24, 25, 32, 35, 41]
total_emp_age = 0for age in emp_age:
    total_emp_age += age
print(total_emp_age)

emp_sno = [1, 3, 5, 7, 8]
sum_emp_sno = [i+2 for i in emp_sno]
print(sum_emp_sno)

get_emp_sno = [i+2 for i in emp_sno if i<5]
print(get_emp_sno)

for i in range(1,10):
    if(i==5):
        break    print(i)
print('Done')

File System in Python - Python supports a number of formats for file reading and writing.
                                   
- In order to open a file, use open() method specifying file name and mode of opening(read,write,append..etc)
- Open() returns a file handle.
- handle = Open(filename, mode)
- Finally we need to close the file using close() method.(Otherwise other programs might not be able to access the file.

File Exercises:
# Files - Open & Close Excerciseemp_file = open('C:/Python_Testing/file1.txt','r')
#for line in emp_file:#    print(line)print('File name is ',emp_file.name)
test = emp_file.read()
print(test)
emp_file_out = open('C:/Python_Testing/file2.txt', 'w')
#emp_file_out.seek(0,0)print(emp_file_out.write(test))
emp_file_out.close()
emp_file.close()
We can access any type of file in Python. Example convert SAS file to Text file and access the data. Python has all specific packages for all type of files.
import sas7bdat
from sas7bdat import *
# To convert a SAS file to a text file.
data = SAS7BDAT('C:/Python_Testing/e.sas7bdat')
data.convertFile('C:/Python_Testing/e_data.txt', '\t')


Functions - Reusable piece of software.
Block of statements
- That accepts some arguments. Function accepts any number of arguments.
- Perform some functionality & provide the output.
- define using def keyword.
- Scope of the variables defined inside the function is Local. These variables can't be used outside of a function.

Functions Exercises:
# Function Excercisedef sayHello():
    print('Hello World !!!')
sayHello()

def printMax(a,b):
    if(a > b):
        print(a, 'is Max value')
    elif(a == b):
        print(a, 'is equal to', b)
    else:
        print(b, 'is Max value')
printMax(5,3)

def hello(message, times=1):
    print(message * times)
hello('Welcome', 5)
hello('Hi')

def varfunc(a, b=5, c=10):
    print('a is', a, 'and b is', b, 'and c is', c)
varfunc(3,4)
varfunc(1,2,3)
varfunc(10,c=12)
varfunc(c=15,a=10)

x=20def varscopefunc(x):
    print('x value is ', x)
    x = 2    print('x local value changed to ', x)
varscopefunc(x)
print('x value is not changed', x)

Modules - Functions can be re/used with in the same program. If you want to use functions outside of the programs that can be achieved using Modules.

- Module is nothing but a package of functions.
- Modules can be imported in other programs & functions contained in those Modules can be used.
- Module create - create a .py file with functions defined in that file.
- Python has huge list of Modules, which are pre-defined. We just need to re/use those modules by using import.

Modules Exercises:
def sayHi():
    print('This is mymodule function')
    return# Factorial Programdef factorial(number):
    product = 1    for i in range(number):
        product = product * (i + 1)
    print(product)
    return product
# C:\Users\khasi\AppData\Local\Programs\Python\Python37-32\Lib\mymodule.py


from mymodule import *
print(sayHi())

numb = int(input('Enter a non-negative number : '))
num_factorial = factorial(numb)
print(num_factorial)

Main function:
import sys
# print('First Line')def main():
    print('Hello World!!!', sys.argv[0])
if __name__ == '__main__':
    main()

Exception Handling - Program will terminate abruptly if you don't handle exceptions at run time. Exceptions are handled in Python using Try-Except & Try-Except-Finally blocks.

Finally is an optional block. Finally block will be executed with or with out exception in the program.

Exception Handling Exercises:
def avg(numlist):
    ''' raise TypeError or ZeroDivisionError Exceptions.'''    sum=0    for num in numlist:
        sum = sum + num
    return float(sum)/len(numlist)

def avgReport(numlist):
        try:
            m = avg(numlist)
            print('Avg is =',m)
        except TypeError:
            print('Type Error')
        except ZeroDivisionError:
            print('Zero Division Error')

list1=[10,20,30,40]
list2=[]
list3=[10,20,30,'abc']

avgReport(list1)
print(avgReport(list2))
print(avgReport(list3))


def avg(numlist):
    ''' raise TypeError or ZeroDivisionError Exceptions.'''    sum=0    for num in numlist:
        sum = sum + num
    return float(sum)/len(numlist)

def avgReport(numlist):
        try:
            m = avg(numlist)
            print('Avg is =',m)
        except TypeError:
            print('Type Error')
        except ZeroDivisionError:
            print('Zero Division Error')
        finally:
            print('Finished avg program')

list1=[10,20,30,40]
list2=[]
list3=[10,20,30,'abc']

avgReport(list2)

Tuesday, December 31, 2019

BigData MapReduce Concept

MapReduce

Is a programming Model to process large datasets in parallel.
MapReduce divides the task into subtasks and handles them in parallel.
Input and Output always be a key-value format.

  • Map
  • Reducer

Map - You have to write a program that can produce local(key, value) pairs.
Eg:- If your data(ex. data is x,y,z) is in 3 data nodes.
After Map - you will get a key-value pair. i.e, data is key & count is value.
DataNode1 o/p => (x,3),(y,3),(z,4)
DataNode2 o/p => (x,3),(y,3),(z,4)
DataNode3 o/p => (x,4),(y,4),(z,3)

Shuffle - Single key and all the values of that key are brought together. This will happen automatically.
Eg:- After the Shuffle phase, for each key, the value from all the data nodes will be accumulated.
(x,(3,3,4))
(y,(3,3,4))
(z,(4,4,3))

Reducer - You have to write a program that will read the key from Shuffle & sum the values of a specific key. The output will be a key-value pair.
Eg:- After the Reducer phase, you will get the following output.
(x,10)
(y,10)
(z,10)

BigData Concepts

BigData Concepts

- Big Data is a massive volume of both structured and unstructured data.
- Big Data is characterized by 3V's. i.e., Volume, Velocity and Variety
                 (Volume - How much data is generating)
                 (Velocity - What pace/speed the data is generating)
                 (Variety - Unstructured data. Ex. picture data, video data, log data..)
- Problems of Big Data is :
                 Storage of the large volumes of Data.
                 Processing of the large volumes of Data.
                 Resource Management.

Hadoop

- It is an Infrastructure. It is a software by using which we can solve the Big Data problems.
  • HDFS - Distributed File System - Solve the problem of Storage.
                                         Horizontal scaling
         Distributed DB = Partitioning  (3GB = 3 * 1GB machines) + Replication (Making multiple copies of the same data at different places - Replication factor is 3 - Fault Tolerance) 
  • Map Reduce - Distributed Computing Engine - Solve the problem of Processing
                                         Shared-Nothing Architecture
  • YARN - Cluster/Resource Manager - Solve the problem of Resource Management.
                                                    Scheduling & Coordination

Data Ecosystems in Enterprise:
Hadoop - Is an additional layer, will allows you to process the BigData in the EcoSystem, which was not done before because of many limitations. Hadoop will co-exists with existing technologies/systems and works with them.



Hadoop:
- Inexpensive commodity hardware
- Free Open Source software
- Scalable
- Reliable
- Enable data archival and reporting
- Enable cutting edge analytics

Hadoop Features:
- Designed to store large files (huge data)
- Processing the data sequentially. So it is good for analyzing entire datasets quickly.
- Run large batch processes that may take several hours, keeps running despite partial failure of the cluster.
- Handles Scheduling & Coordination very well.
- Can store and process unstructured data(log files, text, image, audio, vedio...).

Hadoop is not desined For :
- Processing small files.
- Random access of the data retrieval & should not be used to run Transactional applications.
- Interactive querying. Most of the jobs will take at least several minutes to process.
- Not allows modification of data in place. It is a write once, read many times system. New data can be appended to the file.
- Does n't meet ACID standards, 3NF, data quality in the way that relational databases do.
- Hadoop is not a replacement for your existing database systems!

Enterprise = Hadoop(bulk storage for analytics) + RDBMS(for business operations) + NO SQL DB(for run a website)

Hadoop EcoSystem:



Hadoop = HDFS + YARN + MapReduce


MapReduce - A programming framework for parallel processing of data. Hadoop coordinates execution throughout the cluster.

Pig and Hive - These tools allow the user to write data processing programs. Internally those commands are translated into MapReduce jobs.

Cluster - Set of host machines. Its an hardware infrastructure.

YARN - Resource Manager + Node Manager

Resource Manager(Like Project Manager) - Manage all the Resources - One per Cluster. (Installed on Master Machine) - Monitor the resources at regular intervals. If any resource is not working then it will create a backup for that resource.
- Resource Manager not store any kind of data.
Node Manager(Like Project Developer) - Each Machine has one Node Manager. (Installed on all Slave Machines). It will update the status of each machine is working fine to Resource Manager using heart beat. And provide the Resources, means it will update the status(memory, storage space...) information on each machine to Resource Manager. Container is nothing but a Machine.








HDFS - Hadoop Distributed File System on top of UNIX file system.

HDFS get the information of Resources by YARN.
- When user want to store File(1TB) into HDFS, Based on the Resources information, HDFS split the File(1TB) into small pieces/blocks and store in different machines.
- When user wants the File(1TB) back, then HDFS collect all the pieces and combine them and give it back to the user.
- HDFS

HDFS replicate the received File(Ex.,1TB) and store on different machines for backup and Fault tolerance. Default replication factor is 3.

HDFC = Like YARN you have Master/Slave architecture.
NameNode = Master = Installed on only Master machine = Don't store any data = Job is to manage where different blocks of data is stored on slave machines. It has all meta data information(where different data blocks are stored...). Has the information of all the blocks and store location and all meta data/folder structure information.
- Store the Metadata in memory for faster access.
- NameNode has backup called Secondary NameNode in case of Primary NameNode failure.
- Replication factor info store by NameNode but actual replication is done by DataNode. NameNode replicate DataNode blocks in event of failure.
DataNode = Slave =  Installed on all slave machines = Store all the data in different blocks.

HDFC data blocks = Large file into small chunks/blocks = Each chunk/block is 128MB in size.

- Hadoop can easily calculate how many blocks can fit on a Node. These blocks are large enough to read quickly from disk.



Saturday, October 19, 2019

PIG Project work

BigData PIG Project work
  • LOAD
  • FILTER
  • FOREACH ... GENERATE
  • SPLIT
  • GROUP
  • JOIN
  • DESCRIBE
  • EXPLAIN
  • ILLUSTRATE
  • DUMP
> pig -x local
> pig -x local [script]
> pig -x hadoop [script]

Case Study 1:
Movies dataset with 50000 observations. This dataset has 5 columns(Id, Name, Year, Rating, Duration).

1) grunt> mov = load 'Desktop/Basha/Basha2019/PIG_Practicals/movies_data.xls' using PigStorage(',') as (id:int, name:chararray, year:int, rating:float, duration:int);
grunt> describe mov;
mov: {id: int,name: chararray,year: int,rating: float,duration: int}
grunt> dump mov;

2) Movies list which has rating>4.
grunt> mov_ratingfour = filter mov by (float)rating>4.0;

3) List of movies in the file.
grunt> mov_group = group mov all;
grunt> mov_count = foreach mov_group generate COUNT(mov.id);
grunt> dump mov_count;

4) List title and duration from the file & display the list using duration in DESC.
grunt> mov_duration = foreach mov generate name,(double)duration/60;
grunt> mov_notnull = filter mov_duration by $1 is not null;
grunt> mov_duration_order = order mov_notnull by $1 DESC;
grunt> mov_long = LIMIT mov_duration_order 50;
grunt> dump mov_long;

5) Grouping file List using year.
grunt> mov_group_year = group mov by year;
grunt> mov_group_rating = foreach mov_group_year generate group as year, MAX(mov.rating) as highest_rating;
grunt> dump mov_group_rating;

6) JOIN concept in movies dataset.
grunt> mov_join = JOIN mov_group_rating by (year,highest_rating),mov by (year,rating);
grunt> describe mov_join;
mov_join: {mov_group_rating::year: int,mov_group_rating::highest_rating: float,mov::id: int,mov::name: chararray,mov::year: int,mov::rating: float,mov::duration: int}
grunt> mov_best = foreach mov_join generate $0 as year,$3 as title,$1 as rating;
grunt> describe mov_best;
mov_best: {year: int,title: chararray,rating: float}
grunt> dump mov_best;

CaseStudy 2:
We have a demonetization dataset. We will extract the twitter #Demonitisation tweets and we will want to do some kind of sentimental analysis. Like the people are +ve or -ve sentiment about demonetization.

1) Load the demonetization dataset.
grunt> tweet_load = load 'Desktop/Basha/Basha2019/PIG_Practicals/demonitization_tweets.csv' using PigStorage(','); 
2) Extract id, text columns from the tweets.
grunt> tweet_extract = foreach tweet_load generate $0 as id, $1 as text;
grunt> describe tweet_extract;
tweet_extract: {id: bytearray,text: bytearray}

3) Tokenize the text column value.
grunt> tweet_tokens = foreach tweet_extract generate id, text, FLATTEN(TOKENIZE(text)) as word;
grunt> describe tweet_tokens;
tweet_tokens: {id: bytearray,text: bytearray,word: chararray}

4) Using AFINN dictionary, we will define the +ve/-ve words.
grunt> tweet_dictionary = load '/home/cloudera/Desktop/Basha/Basha2019/PIG_Practicals/AFINN.txt' USING PigStorage('\t') as (word:chararray, rating:int);

5) Join the tweet_tokens and tweet_dictionary.
grunt> tweet_join = join tweet_tokens by word left outer, tweet_dictionary by word using 'replicated';
grunt> describe tweet_join;
tweet_join: {tweet_tokens::id: bytearray,tweet_tokens::text: bytearray,tweet_tokens::word: chararray,tweet_dictionary::word: chararray,tweet_dictionary::rating: int}

6) Tweet Rating.
grunt> tweet_rating = foreach tweet_join generate tweet_tokens::id as id,tweet_tokens::text as text, tweet_dictionary::rating as rate;

7) Grouping the word.
grunt> tweet_word_group = group tweet_rating by (id,text);

8) Average of tweet_rating rate value.
grunt> tweet_avg_rate = foreach tweet_word_group generate group,AVG(tweet_rating.rate) as tweet_finalrating;

9) +ve & -ve tweets.
grunt> tweet_positive = filter tweet_avg_rate by tweet_finalrating>=0;
grunt> tweet_negative = filter tweet_avg_rate by tweet_finalrating<0;