← All courses

Training

Course Outline for Data Science using ML and AI in python

Course Outline for Data Science using ML and AI in python

This course has been designed to cater to the needs of learners who are comfortable with python and picking up a specialization of python as an additional concentrated skill. This course will also delve into the various usages of python which in real business case studies. This course aims to alleviate the unnecessary jargon and cover the parts that are indeed needed.

This course shall also enable the user to have the necessary foundation to take their training to a more advanced level should they wish to do so.

Duration

  • 5 days for those who do not know python programming (but know basics of programming in general)
  • 4 days for those who know basic python programming

Learning outcome:

  1. Overview of python basics.
  2. IDEs of python.
  3. Understand the basic syntax structures.
  4. Get updated on recent syntax changes in python 3.8 and 3.9 respectively.
  5. Advanced Data handling in python.
  6. Working with external libraries in python.
  7. Working with files and encoding.
  8. Understand different forms of data acquisition.
  9. Data Extraction and conversion.
  10. Learn statistical analytics of data.
  11. Use Python to get the basics of Data Analytics.
  12. Use simple scripting to present data.
  13. Have the ability to visualize data and manipulate them using simple python.
  14. Understand how data science delves into Machine Learning.
  15. Machine learning and coding in python using Scikit Learn
  16. Machine learning and coding in python usingXGBoost
  17. Using machine learning for predictions and analytics.
  18. Data Acquisition and scraping.

Prerequisite

  • Basics of computer programming
  • Basic ability to program in python
  • Access to Google Colab OR
  • If the user wants to use python locally, they will need administrative access to the OS as well as installation of Anaconda.
  • High Speed internet connection (min. 1mbps)
  • Basic understanding of how the filing system works

Course Content:

Statistics and Probability and tools to use for Data Analytics

  1. Python Revision
  2. A Crash Course in matplotlib.
  3. Advanced Visualization with Seaborn.
  4. Covariance and Correlation.
  5. Conditional Probability.
  6. Bayes' Theorem.

Real hands on Data Analytics

  1. Basic Machine Learning (using Python).
  2. Supervised vs. Unsupervised Learning, and Train/Test.
  3. Using Train/Test to Prevent Overfitting a Polynomial Regression.
  4. Bayesian Methods: Concepts.
  5. Practice with:
  6. Implementing a Spam Classifier with Naive Bayes.
  7. K-Means Clustering.
  8. Entropy
  9. Decision TreesAssessment,

Dealing with Real-World Data

  1. Bias/Variance Tradeoff
  2. K-Fold Cross-Validation to avoid overfitting
  3. Data Cleaning and Normalization
  4. Cleaning web log data
  5. Normalizing numerical data
  6. Detecting outliers
  7. Feature Engineering
  8. Imputation Techniques for Missing Data
  9. Handling Unbalanced Data: Oversampling,
  10. Undersampling, and SMOTE
  11. Binning, Transforming, Encoding, Scaling, and Shuffling

Machine Learning on Big Data

  1. Introducing MLLib
  2. Introduction to Decision Trees in Spark
  3. K-Means Clustering in Spark
  4. TF / IDF
  5. Searching Wikipedia with Spark
  6. Using the Spark with python brief

Day 5: Deployment and ML in the Real World

  1. Deploying Models to Real-Time Systems
  2. Testing Concepts
  3. T-Tests and P-Values
  4. Hands-on With T-Tests
  5. Determining How Long to Run an Experiment
  6. A/B Test Gotchas

Practical, connected learning

My wider training approach brings hands-on implementation and systems thinking together, connecting technology with real operational needs.