← All courses

Training

Course outline for Cloudera CDP Generalist Exam

Course outline for Cloudera CDP Generalist Exam

This Course prepares you with a broad knowledge of the Cloudera CDP platform. The previously offered examinations (Such as CCA-159 or CCA-175 are no longer offered by Cloudera).

Unlike other CDP Certification Program role-based exams, this exam is applicable to multiple roles:

  • Administrator
  • Developer
  • Data analyst
  • Data engineer
  • Data scientist
  • System architect

Whether an experienced professional, or just starting an enterprise data career, this exam allows candidates to demonstrate their broad understanding of CDP.

Course Outcome

By the end of this course, the learner may be expected to comprehend and utilize:

  • HDFS
  • Ozone
  • Hive
  • Hue
  • YARN
  • Spark
  • Impala
  • Oozie
  • Kafka
  • NiFi
  • HBase
  • Phoenix
  • Kudu
  • Shared Data Experience (SDX)
  • CDP Public integration with cloud SSO
  • CDP Private Cloud integration with LDAP, Kerberos
  • CDP Private Cloud Base HDFS transparent encryption
  • CDP Public Cloud security features of cloud providers storage security
  • Cloudera navigator encrypt
  • SSL/TLS implementation
  • Kerberos authentication
  • Cloudera Analytics
  • CDP Public cloud on major cloud infrastructure providers
  • Workload XM
  • Replication Manager

Please Note: The outcome listed above is a guideline and not a guarantee. Every learner is unique and so is their learning ability. Variations in the ability to comprehend and utilize shall differ from pupil to pupil.

Prerequisites

This is a very intensive course and is not meant for an absolute novice to computing. Learners must meet the minimum requirement listed below:

  • Access to PC [either one of these OS: MS Windows / Mac / Linux (ubuntu)]
  • Virtual Machine (software)
  • The ability to remote into a server via SSH
  • PowerShell (with admin rights) or Terminal (for MacOS and Linux)
  • Programming (any language)
  • Basic linux command structures
  • Basic windows CLI commands
  • Understanding of server architecture
  • Basic systems understanding
  • Basic SQL
  • Experience connecting with DB via programming
  • Experience and understanding of filing system
  • Exposure to LDAP
  • Access to commercial cloud platforms:
    • Digital Ocean
    • AWS
    • Azure
    • GCP
  • High Speed Internet connection (minimum 3Mbps per station)
  • Webcam
  • Microphone
  • Dual screens
  • Ability to comprehend the English language (the examination is only available in the English language)
  • Minimum 99% attendance

Course Outline

The course will be conducted over a period of 20 full days. In the event that the learner is a bit rusty on some of the above mentioned prerequisites, there will be some revisiting of the topics as a quick refresher. The course is designed to have perioding assessments and practical hands-on for maximum learning and retention.

The overall contents of the course shall be as follows:

  1. Basics of CDP
    1. Introduction
    2. Introducing the Enterprise Data Cloud
      1. The Evolution of CDP
      2. Characteristics of an Enterprise Data Cloud
      3. From the Edge to AI: An End-to-End Use Case
      4. Assessment
    3. Cloudera Data Platform Overview
      1. Assessment
    4. Workload and Data Management
      1. The Role of an Administrator
      2. SDX
      3. Managing Resources and Costs
      4. Workload Isolation
      5. Data Migration and Replication
      6. Assessment
    5. Data in Motion
      1. The Role of a Data Engineer
      2. Data in Motion Use Cases
      3. Cloudera DataFlow
      4. Assessment
    6. Data Warehousing and Analytics
      1. The Role of a Data Analyst
      2. Data Warehouse and Analytics Use Cases
      3. Operational and Analytic Database Capabilities
      4. CDP Data Warehouse Experience
      5. Assessment
    7. Data Science and Machine Learning
      1. The Role of a Data Scientist
      2. Machine Learning Use Cases
      3. Cloudera Data Science Workbench (CDSW)
      4. Cloudera Machine Learning
      5. Fast Forward Labs
      6. Assessment
    8. Security and Data Governance
      1. The Role of a Data Steward
      2. Data Catalog
      3. Controlling and Auditing Data Access
      4. Assessment
    9. Practical
      1. Challenges and Opportunities
      2. Assessment
  2. CDP Private Cloud Fundamentals
    1. The Enterprise Data Cloud Vision
      1. The Enterprise Data Cloud Vision
      2. Characteristics of the Enterprise Data Cloud
      3. Data Lifecycle
    2. Cloudera Data Platform Overview
      1. Cloudera Data Platform: Recap
      2. How to Eliminate Shadow IT
      3. CDP Public Cloud
      4. CDP Private Cloud Base
    3. Introducing CDP Private Cloud
      1. CDP Product Overview
      2. CDP Private Cloud Plus
    4. CDP Private Cloud Architecture
      1. Important Trends
      2. Traditional Cluster Architecture (Bare Metal)
      3. Limitations of the Traditional Cluster Architecture
      4. Key Aspects of the Cloud-Native Architecture
      5. What is Kubernetes?
      6. Kubernetes Overview
      7. Comparing CDP Public and Private Cloud
      8. CDP Private Cloud Architecture
    5. Installation
      1. Installation Requirements
      2. Installation Demo
      3. Essential Points
    6. Commercial Installations:
      1. Digital Ocean
      2. AWS
      3. Azure
      4. GCP
  3. Programming for Big Data
    1. Python
    2. Scala
    3. Java
  4. Kafka
    1. Kafka Theory
      1. Broker
      2. Cluster
      3. Zookeeper
      4. Ensemble
      5. Multiple clusters
      6. Ports
      7. Topics
      8. Message structure
      9. Partitions and distribution
      10. Producer and Consumer architecture
    2. Kafka deployment
      1. Operating Ubuntu Server
      2. Setting up IDE to work remotely with ubuntu from Windows OS
      3. Installing required environments
      4. INstalling kafka from scala binaries
    3. Using kafka to send and receive messages
      1. Send messages using the console
      2. Consume messages using the console
      3. Running multiple consumers and multiple producers
    4. Using Kafka with multiple partitions
      1. Remote brokers
      2. Creating multiple partitions
      3. Accessing specific partitions
      4. Accessing specific offsets
      5. Exploring and understanding the filing system of messages
    5. Kafka Cluster with multiple brokers
      1. Configurations and settings
      2. Launching and managing
      3. Using Zookeeper directly for info access on brokers
      4. Multiple partition topics in clusters
      5. Logs
      6. Producing and Consuming inter-cluster
      7. Simulating break-downs
    6. Multiple brokers with Replication
      1. Creation and launch
      2. Actual replication
      3. Logs and details exploration
      4. Production and consumption
      5. Storage management
      6. Break-down simulation
      7. Manage and recover from break-downs
    7. Consumer Groups
      1. Using the default consumer groups
      2. Multiple consumer groups
      3. Idle consumer groups
    8. Assessment
    9. Performance and Testing
      1. Basic performance test
      2. Tweaking the parameters
      3. Consumer performance
      4. Nonzero LAG values for consumers
  5. Hadoop and related technologies
    1. Introduction
      1. Big Data Case Studies
      2. Brief History of Hadoop
      3. Need for Hadoop
      4. Hadoop Architecture
      5. RDBMS vs Hadoop
      6. Vendor Comparison
      7. Hardware Recommendations
      8. Hadoop Installation
    2. HDFS
      1. HDFS Basics
      2. HDFS Architecture
      3. Data Read and Write Process
      4. HDFS Permissions
      5. Data Replication
      6. HDFS Accessibility
      7. HDFS Filesystem Operations
      8. HDFS Interfaces
      9. Heartbeats
      10. Rack Awareness
      11. distcp
    3. MapReduce
      1. MapReduce Basics
      2. MapReduce Workflow
      3. MapReduce Framework
      4. Hadoop Data Types
      5. MapReduce Internals
      6. Job Formats
      7. Debugging and Profiling
      8. Distributed Cache
      9. Combiner Functions
      10. MapReduce Streaming
      11. MapReduce Counters, Sorting and Joins
    4. YARN
      1. YARN Infrastructure
      2. YARN ResourceManager
      3. YARN ApplicationMaster
      4. YARN NodeManager
      5. YARN Container
    5. Apache Pig
      1. Pig Architecture
      2. Pig Installation and Modes
      3. Grunt and Pig Script
      4. Pig Latin Commands
      5. UDF and Data Processing Operator
    6. HBase
      1. HBase Architecture
      2. HBase Installation
      3. HBase Configuration
      4. HBase Schema Design
      5. HBase Commands
      6. MapReduce Integration
      7. HBase Security
    7. Sqoop
    8. Flume
    9. Hive
      1. Hive Architecture
      2. Hive shell
      3. Hive Data Types
      4. HiveQL
    10. Hadoop Workflow
    11. Hadoop Cluster Management
      1. Cluster Planning
      2. Cluster Installation and Configuration
      3. Cluster Testing
      4. Cluster Benchmarking
      5. Cluster Monitoring
    12. Hadoop Administration
      1. dfsadmin, fsck and balancer
      2. Hadoop Logging
      3. Hadoop Data Backup
      4. Addition and removal of nodes
    13. Hadoop Security
      1. Authentication
      2. Data Confidentiality
      3. Configuration
    14. NextGen Hadoop
      1. HDFS HA
      2. HDFS Federation
  6. Spark
    1. Spark basics and RDD interface
    2. Hands in installation and setup
    3. Using libraries and resources
    4. Exercises
    5. Assessment
    6. SparkSQL, DataFrames and Datasets
    7. Advanced Spark
    8. Accumulators and BFS in Spark
    9. Clustering
    10. Machine Learning with Spark
    11. Streaming
    12. GraphUX
    13. Programming with Pig
    14. Spark and RDD revisit
    15. Hive
    16. RDBMS with Hadoop
    17. Flat-DB with Hadoop
    18. Querying
    19. Apache Storm
    20. Real World problems exercise
  7. Security
    1. Course Introduction
      1. Security Overview
    2. Setup
      1. Zookeeper Setup
      2. Hands-On: Setup Zookeeper Service
      3. Test
      4. Kafka Setup: Producer / Consumer test
    3. SSL Encryption in Kafka
      1. The need for SSL Encryption
      2. SSL - hands on
      3. Hands-On: Creating a Certificate Authority (CA)
      4. Hands-On: SSL Setup in Apache Suite
      5. Hands-On: SSL Setup for Clients
      6. SSL Encryption in Apache Suite
    4. SSL Authentication in Apache Suite
      1. What is SSL Authentication?
      2. Hands-On: SSL Authentication
    5. SASL Authentication - Kerberos / GSSAPI in Kafka
      1. What is SASL in Kafka?
      2. What is Kerberos?
      3. Hands-On Kerberos - Part 1: Setup
      4. Hands-On Kerberos - Part 2: Principals & Keytabs
      5. Hands-On Kerberos - Part 3: Kafka Configuration
      6. Hands-On Kerberos - Part 4: Client Configuration
    6. Authorization in Kafka
      1. ACLs in Kafka
      2. Hands-On: ACL demo
    7. Zookeeper Security
      1. Introduction
      2. Zookeeper Create Principal
      3. Zookeeper Configure Kerberos
      4. Hands-On: ZNode General
      5. Zookeeper Authorisation Config
      6. Hands-On: Zookeeper SuperUser
      7. Zookeeper Security Migration Tool and Summary
    8. Cluster Security
  8. Analytics
    1. Using Cloudera QuickStart VM
      1. Installation
      2. Virtualization
      3. SSH and remoting
    2. Overview of Big Data ecosystem
      1. Overview of Distributions and Management Tools such as Ambari
      2. Properties and Properties Files of Big Data Tools - General Guidelines
      3. Hadoop Distributed File System - Revision
      4. Distributed Computing using YARN and Map Reduce 2 - revision
      5. Submitting Map Reduce Job in YARN Framework
      6. Determining Number of Mappers and Reducers
      7. Understanding YARN and Map Reduce Configuration Properties
      8. Reviewing and Overriding Map Reduce Job Run Time Properties
      9. Map Reduce Job Counters
      10. Overview of Hive
      11. Databases in Big Data and Query Engines
      12. Overview of Data Ingestion Tools in BigData
    3. Overview of HDFS Commands - revision
      1. Revision of "hadoop fs" or "hdfs dfs" command
      2. Listing Files in HDFS
      3. User Spaces or Home Directories in HDFS
      4. Creating Directory in HDFS
      5. Copying Files and Directories into HDFS
      6. File and Directory Permissions Overview
      7. Getting Files and Directories from HDFS
      8. Previewing Text Files in HDFS - cat and tail
      9. Copying or Moving Files from one HDFS location to other HDFS location
      10. Understanding Size of the File System and Data Sets - using df and du
      11. Overview of Block Size and Replication Factor
      12. Getting metadata of files using "hdfs” and “fsck"
    4. Apache Hive - for analytics
      1. Revisiting Hive Language Manual
      2. Launching and Using Hive CLI
      3. Overview of Hive Properties - SET and .hiverc
      4. Hive CLI History and .hiverc
      5. Running HDFS Commands using Hive CLI
      6. Understanding Warehouse Directory
      7. Creating Database in Hive and Switching to the Database
      8. Creating First Table in Hive and list the tables
      9. Retrieve metadata of HiveRole of Hive Metastore
      10. Beeline - Alternative to Hive CLI
      11. Default Delimiters in Hive Tables using Text File Format
      12. Revisiting File Formats - STORED AS Clause
      13. Truncating and Dropping tables in Hive
    5. Partitioning and Bucketing
      1. Partitioning and Bucketing in Hive
      2. Creating Tables using orc File Format - order_items
      3. Inserting Data into order_items using stage table
      4. LOAD Command
      5. Creating Partitioned Tables in Hive -
      6. Adding Partitions to Tables in Hive
      7. Loading into Partitions in Hive Tables
      8. Inserting Data into Partitions in Hive
      9. Creating Bucketed Tables
      10. Bucketing with Sorting
      11. Overview of ACID Transactions in Hive
      12. Updating and Deleting data in Hive Bucketed Tables
    6. Overview of Functions
      1. Functions
      2. Validating Functions
      3. String Manipulation
      4. Date Manipulation
      5. Overview of Numeric Functions
      6. Type Cast Functions for Data Type Conversion
      7. Handling null values using nvl
    7. Writing Queries
      1. Overview of SQL
      2. Reviewing Logs for Hive Queries
      3. Projecting Data
      4. Projecting DISTINCT Values
      5. Filtering Data
      6. Boolean Operations
      7. Basic Aggregations using Aggregate Functions
      8. Sorting Data with in groups
      9. Overview of CLUSTER BY
    8. Joins and Set Operations
      1. Nested Sub Queries
      2. Joins
      3. Joining Multiple Tables in Hive
      4. Cartesian between two data sets
      5. Union between two Data Sets
      6. Troubleshooting
    9. Analytics and Windowing Functions
      1. Hands on - HR Database
      2. Overview of Analytics Functions
      3. Windowing Functions
      4. Performing Aggregations
      5. CRUD
      6. Applying rank Function
      7. Applying dense_rank Function
      8. Applying row_number Function
      9. Understanding Order of Execution
      10. Filtering data using fields derived using analytics or windowing functions
    10. Running Queries using Impala
      1. Introduction to Impala
      2. Role of Impala Daemons
      3. Impala State Store and Catalog Server
      4. Overview of impala-shell
      5. Relationship between Hive and Impala
      6. Loading and Inserting Data into Impala Tables
      7. Running Queries using Impala Shell
      8. Reviewing Logs of Impala Queries
      9. Synching Hive Metadata with Impala -
      10. Running Scripts using Impala Shell
      11. Develop and run Impala Script
    11. Apache Sqoop
      1. Introduction to Sqoop
      2. Validate Source Database
      3. Getting help of Sqoop using Command Line
      4. Overview of Sqoop User Guide
      5. Validate Sqoop
      6. List tables in MySQL using "sqoop list-tables"
      7. Run Queries in MySQL using "sqoop eval"
      8. Understanding Logs in Sqoop
      9. Redirecting Sqoop Logs into files
    12. Apache Sqoop - Importing Data into HDFS
      1. Overview of Sqoop Import Command
      2. Perform Sqoop Import
      3. Managing HDFS
      4. Reviewing logs of Sqoop Import
      5. Output Files
      6. Validating avro Files using avro-tools
      7. Using Compression
      8. Sqoop Import
    13. Apache Sqoop - Importing Data into Hive Tables
      1. Sqoop Import
      2. Understanding Execution Flow
    14. Apache Sqoop - Exporting Data from HDFS to RDBMS
      1. Prepare data for Export
      2. Sqoop Export
    15. Apache Sqoop - Incremental Imports and Jobs
      1. Overview of Sqoop Jobs
      2. Creating Sqoop Job
      3. Running Sqoop Job
      4. Overview of Incremental Imports
      5. Incremental Import
  9. Practice runs
    1. Assessment mock - 1
    2. Assessment mock - 2

Exam Details

  • Exam Number: CDP-0011
  • Number of questions: 60
  • Duration: 90 minutes
  • Pass Score: unpublished

Cloudera does not publish exam pass scores. Candidates should not be trying to achieve any particular score. Rather they should be aiming for the highest score possible.

  • Delivery: online, proctored
  • Real-time communication components
  • Functional and enabled for real-time communication via Zoom with the exam proctor.
    • A browser with pop-up blocker disabled
    • A built-in or external webcam and microphone
    • Internet speed must be at least 2 Mbps download and 2 Mbps upload.
    • Online proctoring uses Zoom for real-time communication
    • Once connected to an exam, wait for guidance from the proctor about logging in to Zoom.
  • Allowed resources: none.

You may not use reference materials, white papers, user guides, or any other resources during your exam.

Test Cost: USD 300

Practical, connected learning

My wider training approach brings hands-on implementation and systems thinking together, connecting technology with real operational needs.