Part 1:
Learning outcome:
- Install and run Apache Spark on a desktop computer or on a cluster
- Frame big data analysis problems as Spark problems
- Implement iterative algorithms such as breadth-first-search using Spark
- Use DataFrames and Structured Streaming in Spark 3
- Use Spark's Resilient Distributed Datasets to process and analyze large data sets across many CPU's
- Understand how Spark Streaming lets your process continuous streams of data in real time
- Share information between nodes on a Spark cluster using broadcast variables and accumulators
- Analyze non-relational data using HBase, Cassandra, and MongoDB
- Publish data to your Hadoop cluster using Kafka, Sqoop, and Flume
- Use Pig and Spark to create scripts to process data on a Hadoop cluster in more complex ways.
- Analyze relational data using Hive and MySQL
- Understand how Hadoop clusters are managed by YARN, Tez, Mesos, Zookeeper, Zeppelin, Hue, and Oozie (optional)
- Consume streaming data using Spark Streaming, Flink, and Storm (optional)
Course Content:
Day 1:
- Spark basics and RDD interface
- Hands in installation and setup
- Using libraries and resources
- Exercises
- Assessment
Day 2:
- SparkSQL, DataFrames and Datasets
- Advanced Spark
- Accumulators and BFS in Spark
Day 3
- Clustering
- Machine Learning with Spark
- Streaming
- GraphUX
Day 4
- Programming with Pig
- Spark and RDD revisit
- Hive
- RDBMS with Hadoop
- Flat-DB with Hadoop
Day 5
- Querying
- Apache Storm
- Real World problems exercise
Practical, connected learning
My wider training approach brings hands-on implementation and systems thinking together, connecting technology with real operational needs.