FA-0751Data & AnalyticsDevOps, Cloud & InfrastructureSoftware Development

Apache Spark and Hadoop Ecosystem Processing

Batch analytics, streaming, graph processing and data integration

Introduction

Why this course

Build practical Spark data-processing workflows and connect them to the wider Hadoop ecosystem. The programme combines guided installation, coding exercises and a final applied problem across five days.

Core Spark exercises receive hands-on attention. The broader ingestion, database and ecosystem landscape is explored through selected demonstrations and comparisons, with legacy technologies clearly distinguished from current deployment options.

Learning outcomes

Learning outcomes

  • Set up a compatible local Spark environment and explain how an application runs on a cluster.
  • Frame a data-analysis problem as distributed transformations and actions using RDDs, Spark SQL and DataFrames.
  • Use broadcast variables and accumulators appropriately, and work through a breadth-first-search example.
  • Explore clustering, machine-learning and graph-processing examples using the appropriate Spark APIs.
  • Build a guided Structured Streaming example and explain the older Spark Streaming/DStreams model.
  • Compare relational, flat-file and non-relational integration patterns, including Hive, MySQL, HBase, Cassandra and MongoDB.
  • Explain the roles of ingestion, execution, coordination, notebook and orchestration tools without treating them all as cluster managers.
Prerequisites

Prerequisites

Basic programming, SQL and command-line familiarity. A compatible Spark runtime and prepared datasets are required; cluster and external-system demonstrations use a prepared teaching environment.

Training outline

5 modules

·
01Day 1 — Spark foundations and setup3 topics
  • Spark execution concepts and the RDD interface: transformations, actions and distributed processing.
  • Guided local installation, libraries, resources and a first data-processing workflow.
  • Practice exercises and a foundation assessment.
02Day 2 — Structured data and iterative processing4 topics
  • Spark SQL, DataFrames and the Dataset API in languages that support it.
  • Advanced processing patterns and an introduction to partitioning and reuse of intermediate results.
  • Broadcast variables and accumulators: appropriate uses and limitations.
  • Guided iterative breadth-first-search exercise.
03Day 3 — Clustering, machine learning, streaming and graphs4 topics
  • Distributed execution and clustering concepts; a selected Spark machine-learning example.
  • Structured Streaming with DataFrames, checkpointing and a prepared streaming input.
  • Older DStreams-based Spark Streaming as a legacy comparison; latency and delivery behaviour depend on mode and connectors.
  • GraphX graph-processing concepts and a selected Scala example.
04Day 4 — Hadoop data processing and integration5 topics
  • Pig scripting as a legacy Hadoop processing approach; revisit equivalent Spark and RDD patterns.
  • Hive querying, relational integration with MySQL and flat-file data processing.
  • HBase, Cassandra and MongoDB integration concepts with selected prepared examples.
  • Kafka ingestion and the historical roles of Flume and retired Sqoop; choose maintained connectors for current exercises.
  • Distinguish YARN resource management, Tez execution, ZooKeeper coordination and Zeppelin/Hue user tools. Spark deployment options include standalone, YARN and Kubernetes; Mesos and retired Oozie are historical ecosystem context.
05Day 5 — Querying and applied processing3 topics
  • Query integrated datasets and compare batch and streaming problem requirements.
  • Apache Storm concepts and a guided example or prepared demonstration; optional comparison with Flink and Spark streaming.
  • Apply selected techniques to a real-world-style dataset, review results and discuss trade-offs.

A programme built around your team.

Share your training goals and requirements.

Apache Spark and Hadoop Ecosystem Processing
FA-0751

Share your requirements for this programme.

Training enquiry

Apache Spark and Hadoop Ecosystem Processing