← All courses

Training

R Advanced

R Advanced

Data Science and Machine Learning

Data science and analytics is an exciting discipline that allows you to turn raw data into understanding, insight, and knowledge. The goal of this course is to help you learn the most important tools in R that will allow you to do data science and use machine learning for predictive analytics. This course is designed to give you the skills needed to tackle a wide variety of data science challenges using the best parts of R.

Learning Outcome:

After the completion of this course successfully, the learner shall be exposed and will have trained on the following:

  • Data pre-processing
  • Data cleaning
  • Data Cleaning
  • Data wrangling
  • Data Sourcing
  • Advanced predictive analytics
  • Machine Learning with R
  • Various models of ML
  • Predictive Analytics with R
  • Application of techniques to real-world data though with real-world examples
  • Rcpp, ggplot2, and dplyr
  • Advanced visualization techniques

Prerequisites

In order to make the most out of this course, the learner needs to qualify with the following few criterion:

  • Fundamentals of programming in R
  • Ability to source and handle data
  • Basic knowledge of linear mathematics
  • Knowledge in non-linear mathematics will be an added advantage but not required
  • Understanding of filing system in OS of choice

Course Outline:

Duration: 2 days

Intensity: High

Difficulty: Medium

  1. R Refresher
    1. The Distribution of Data
    2. Univariate Data
    3. Frequency Distribution
    4. Central Tendency
    5. Spread
    6. Probability Distribution
    7. Visualization
    8. Exercises
  2. Relationship between data
    1. Multivariate data
    2. Relationships between a categorical and continuous variable
    3. Relationships between two categorical variables
    4. The relationship between two continuous variables
    5. Visualization methods
    6. Exercises
  3. Probability
    1. Basic probability
    2. Sampling from distributions
    3. The normal distribution
    4. Exercises
    5. Using Data To Reason About The World
    6. Estimating means
    7. The sampling distribution
    8. Interval estimation
    9. Smaller samples
    10. Exercises
  4. Testing Hypotheses
    1. The null hypothesis significance testing framework
    2. Testing the mean of one sample
    3. Testing two means
    4. Testing more than two means
    5. Testing independence of proportions
    6. What if my assumptions are unfounded?
    7. Exercises
  5. Bayesian Methods
    1. The big idea behind Bayesian analysis
    2. Choosing a prior
    3. Using MCMC
    4. Using JAGS and runjags
    5. Fitting distributions the Bayesian way
    6. The Bayesian independent samples t-test
    7. Exercises
  6. The Bootstrap
    1. Performing the bootstrap in R
    2. Confidence intervals
    3. A one-sample test of means
    4. Bootstrapping statistics other than the mean
    5. Exercises
  7. Sources of Data
    1. Relational databases
    2. Using JSON
    3. XML
    4. Other data formats
    5. Online repositories
    6. Exercises
  8. Dealing with Missing Data
    1. Analysis with missing data
    2. Visualizing missing data
    3. Types of missing data
    4. Unsophisticated methods for dealing with missing data
    5. So how do we come up with the imputed values?
    6. Exercises
  9. Dealing with Messy Data
    1. Checking unsanitized data
    2. Regular expressions
    3. Other tools for messy data
    4. Exercises
  10. Dealing with Large Data
    1. Wait to optimize
    2. Using a bigger and faster machine
    3. Be smart about your code
    4. Using optimized packages
    5. Using another R implementation
    6. Using parallelization
    7. Using Rcpp
    8. Being smarter about your code
    9. Exercises

Practical, connected learning

My wider training approach brings hands-on implementation and systems thinking together, connecting technology with real operational needs.