Course outline on Big Data for beginners using Hadoop
This 7 day course jumps into the world of Hadoop ecosystem components and its tools in a simplified manner, and provides you with the skills to utilize them effectively for faster and effective development of Big Data projects.
Course Outcome
Every learner has different learning patterns and absorption rates. As such, it is virtually impossible to gauge an exact learning outcome of each participant. However, it may be safely assumed that by the end of this course, each learner shall be able to:
- Understand how Hadoop clusters are managed by YARN, Tez, Zookeeper, Zeppelin,
- Install and run Apache Hadoop and its components on a cluster
- Frame big data analysis problems
- Implement iterative algorithms such as breadth-first-search
- Use DataFrames and Structured Streaming
- Use Spark's Resilient Distributed Datasets to process and analyze large data sets across many CPU's
- Understand how Spark Streaming lets your process continuous streams of data in real time
- Share information between nodes
- Analyze non-relational data using HBase, Cassandra, and MongoDB
- Publish data to your Hadoop cluster using either Kafka, Sqoop, or Flume
- Use Spark to create scripts to process data on a Hadoop cluster in more complex ways.
- Analyze relational data using Hive
- Python Programming
- Python for Spark
- Use the Linux operating system for setup of clusters
- Understand basic security systems pertaining to Big Data
Prerequisites
It is expected that the pupil shall have the following at hand when participating in the course:
- Ability to comprehend the English Language (basic)
- Access to High Speed Internet
- Access to Linux based operating system (at least 3 units)
- Basic understanding of OS
- Basic understanding and experience using the internet
- Basic understanding of filing systems
- High School Mathematics
Course Outline
- Prerequisites
- Programming
- Python
- Introduction
- IDE
- Variables
- Data Types
- Conditions
- Loops
- Functions
- Recursion
- File Systems
- Database
- Java
- Introduction
- IDE
- Variables
- Data Types
- Conditions
- Loops
- Functions
- OOP
- Deployment
- Python
- OS and Networking
- Terminal
- SSH
- Linux Commands
- Linux Scripts
- Linux file system
- Keys
- Bash and profiles
- Programming
- Hadoop
- Introduction
- Background in Big Data
- Overview of components
- Installation and multi-node cluster setup
- MapReduce
- Ambari
- Yarn
- Pig
- Installation
- Pig Latin
- Use Case
- Hive
- Installation
- Queries
- Hive Partitioning, Buckets, UDFs, and SerDes
- Kafka
- Installation
- Single node setup
- Multi-node setup
- Multi-node cluster with replication
- Programming for Kafka
- Python
- Java
- Spark
- Spark Execution Model Architecture
- Programming in Spark with python
- Structured API foundation
- Dataframes
- Sinks
- Transformations
- Aggregations
- Joins
- Hadoop Security
- Introduction to Kerberos
- Kerberos on Hadoop
- Kerberos Terminology
- Demo: Enabling Kerberos
Practical, connected learning
My wider training approach brings hands-on implementation and systems thinking together, connecting technology with real operational needs.