Real-Time Data Processing with Apache Flink and Iceberg
Harness the Power of Stream Processing and Advanced Table Formats for Modern Data Architectures in 2 days
Data doesn't sleep, and the companies that win don't either. Apache Flink has changed the game by letting you process massive streams of data the moment it arrives—whether it's flowing non-stop or coming in batches. It's fast, it's reliable, and it's what the most innovative teams are using to stay ahead.
Enter Apache Iceberg: the missing piece that brings order to data lake chaos. It keeps your massive datasets consistent and manageable, even when you're working at enterprise scale.
Our two-day workshop cuts through the hype and gets your hands dirty with both technologies. You'll walk away knowing how to build actual real-time data systems that solve real problems—the kind that make executives notice and competitors worry.
Learning Outcomes:
By the end of this course, participants will be able to:
- Comprehend the core concepts and architecture of Apache Flink and its role in stream processing.
- Develop and deploy data processing pipelines using Flink's DataStream API.
- Implement stateful stream processing and manage event time semantics effectively.
- Integrate Apache Flink with Apache Iceberg to manage large analytic tables efficiently.
- Utilize Flink SQL for querying and manipulating data within Iceberg tables.
- Apply best practices for building resilient and scalable real-time data applications.
Prerequisites:
Participants should have:
- Proficiency in Java or Scala programming languages.
- Fundamental understanding of distributed systems and parallel processing concepts.
- Basic knowledge of data processing frameworks such as Apache Hadoop or Apache Spark.
- Familiarity with SQL and relational database concepts.
Detailed Course Outline:
- Introduction to Stream Processing and Apache Flink
- The Evolution of Data Processing: From Batch to Stream
- Key Concepts in Stream Processing
- Overview of Apache Flink
- Flink's Position in the Big Data Ecosystem
- Core Features and Capabilities
- Use Cases and Industry Applications
- Apache Flink Architecture and Runtime
- Flink's Distributed Architecture
- JobManager and TaskManager Roles
- Execution Model and Dataflow
- Fault Tolerance Mechanisms
- Checkpointing and State Snapshots
- Exactly-Once Semantics
- Deployment Modes
- Standalone Cluster
- YARN and Kubernetes Integrations
- Flink's Distributed Architecture
- Getting Started with Flink's DataStream API
- Setting Up the Development Environment
- Building a Simple Flink Application
- Defining Data Sources and Sinks
- Applying Transformations
- Understanding Streams and Transformations
- Stateless vs. Stateful Transformations
- Keyed Streams and Partitioning
- Stateful Stream Processing
- The Importance of State in Stream Processing
- Managing Operator State and Keyed State
- Implementing Process Functions
- Timers and Event Time Processing
- State Backends and Their Configurations
- Memory State Backend
- RocksDB State Backend
- Time and Windowing in Flink
- Understanding Event Time vs. Processing Time
- Generating and Assigning Timestamps
- Watermarks and Their Role in Event Time Processing
- Window Operators and Functions
- Tumbling Windows
- Sliding Windows
- Session Windows
- Late Data Handling Strategies
- Integrating Apache Flink with Apache Iceberg
- Introduction to Apache Iceberg
- Motivation and Key Features
- Comparison with Traditional Table Formats
- Setting Up Iceberg with Flink
- Configuring Iceberg Catalogs
- Creating and Managing Iceberg Tables
- Performing Data Operations
- Inserting and Deleting Records
- Schema Evolution and Partitioning
- Querying Iceberg Tables with Flink SQL
- Writing and Executing SQL Queries
- Optimizing Query Performance
- Introduction to Apache Iceberg
Practical, connected learning
My wider training approach brings hands-on implementation and systems thinking together, connecting technology with real operational needs.