FA-0655Data & AnalyticsSoftware Development

Practical PySpark Data Processing

DataFrames, SQL and selected analytics workflows

Use PySpark DataFrames and SQL to build a small data pipeline, inspect execution plans and explore selected MLlib and Structured Streaming examples in a prepared environment.

Introduction

Why this course

This two-day intensive course is for Python users with basic SQL and distributed-systems knowledge. It develops practical PySpark data-processing foundations using representative training datasets and a prepared local or managed lab.

DataFrame operations, Spark SQL, schemas, data quality and query-plan interpretation form the core practice. Selected machine-learning and streaming exercises demonstrate how these foundations extend to analytical pipelines. Cluster deployment, advanced tuning and the broader algorithm catalogue are overviews or demonstrations, not complete implementations.

The source's broad project scope is consolidated into a small transformation pipeline, one prepared ML example and one streaming-output demonstration. Participants examine limitations and next steps rather than claiming mastery, production scale, hard real-time behaviour or guaranteed performance gains.

Learning outcomes

Learning outcomes

The course teaches participants to:

  • Explain Spark architecture and the roles of RDDs, DataFrames and SQL in PySpark.
  • Use a prepared environment, SparkSession and a notebook or shell.
  • Load selected CSV, JSON and Parquet data, define schemas and apply transformations.
  • Use joins, aggregations, window functions and SQL queries, with data-quality checks.
  • Inspect query plans and reason about caching, partitioning and UDF costs.
  • Build and evaluate one small DataFrame-based ML pipeline with appropriate data separation.
  • Run a selected Structured Streaming example and explain event time, watermarks, state and checkpoints.
  • Identify cluster, resource, integration and monitoring work needed beyond the training lab.
Prerequisites

Prerequisites

  • Basic knowledge of Python programming.
  • Familiarity with data analysis concepts and libraries like Pandas and NumPy (preferred but not mandatory).
  • Understanding of fundamental concepts in distributed systems.
  • Basic knowledge of SQL and database operations.
  • Experience with data processing or analytics (preferred but not mandatory).
  • A laptop with Python installed (specific setup instructions will be provided before the course).

The training setup must include mutually compatible Spark, Python, Java and notebook/client dependencies, with sample data and any required storage access prepared in advance. Familiarity with model evaluation is helpful for the selected ML demonstration.

Training outline

2 modules

·
01Day 1 — Environment, DataFrames and SQL1 topics

Spark and PySpark foundations

  • Distributed processing and Spark driver/executor concepts; a brief history and ecosystem overview.
  • RDDs and DataFrames; typed Dataset API is a Scala/Java concept, not a separate PySpark API.
  • Python versus other Spark APIs, trade-offs and representative use cases.

Environment and configuration

  • Prepared local setup and cluster-deployment options overview.
  • Spark configuration and memory/resource considerations.
  • PySpark shell, SparkSession and Jupyter-based exploration.

DataFrames and transformations

  • DataFrame versus RDD; schema definition, inference and validation.
  • Load CSV, JSON and Parquet; create DataFrames from Python collections or RDDs where supported.
  • Selection, filtering, projection, grouping, joins and unions.
  • Arrays, maps, structs and selected UDF examples; prefer built-in functions where appropriate.
  • Caching/persistence, partitioning and bucketing concepts with measured trade-offs.

Spark SQL and data quality

  • SparkSession, temporary views and tables.
  • Selected SQL operations and analytical window functions.
  • Catalyst/query-plan inspection and a bounded tuning example.
  • Schema checks and data cleansing within a small transformation-pipeline lab.
02Day 2 — ML Pipelines, Streaming and Integrated Labs1 topics

Selected DataFrame-based MLlib workflows

  • DataFrame-based pyspark.ml versus legacy RDD-based pyspark.mllib.
  • Feature extraction, transformation, scaling and normalisation.
  • Pipeline estimators, transformers and evaluation with suitable training/validation separation.
  • One selected classification or regression exercise; clustering and recommendation overview.
  • CrossValidator, TrainValidationSplit and a small ParamGridBuilder grid; random parameter sampling is a separate approach, not a built-in RandomizedSearchCV equivalent.

Structured Streaming fundamentals

  • Batch versus incremental streaming; legacy DStream API context.
  • Sources and sinks for one prepared streaming example.
  • Stateful processing, event-time windows and counts within windows; count-bounded windows require separate logic rather than the standard time-window API.
  • Late data and watermark limits, checkpoints and output-mode considerations.
  • Monitoring and debugging; latency and delivery guarantees depend on processing mode, sources, sinks and implementation.

Lab consolidation and follow-on work

  • Review the environment and transformation pipeline from Day 1.
  • Build and evaluate the selected ML example using prepared data.
  • Inspect streaming analytical output in a prepared notebook or dashboard view; full dashboard infrastructure is not built during the course.
  • Review correctness, resource use, performance observations and next steps for real datasets.

A programme built around your team.

Share your training goals and requirements.

Practical PySpark Data Processing
FA-0655

Share your requirements for this programme.

Training enquiry