FA-0720Data & AnalyticsSoftware Development

Data Science Developer Foundations and Certification Preparation

Python-based data preparation, statistics and introductory machine learning

Introduction

Why this course

Build a foundation in Python-based data science, from data preparation and visualisation to selected supervised and unsupervised models. The seven-day programme combines guided implementation, statistical interpretation and a review exercise using small datasets.

The original source refers to the Global Tech Council Certified Data Science Developer credential. This training does not award that certification, guarantee employment or imply that every algorithm can be mastered in one week; current assessment terms must be checked with the provider.

Learning outcomes

Learning outcomes

  • Use foundational Python, NumPy and pandas to work with tabular data.
  • Prepare and visualise sample datasets and explain key statistical summaries.
  • Distinguish regression, classification, clustering and dimensionality-reduction goals.
  • Implement selected model examples and evaluate them with suitable data separation and metrics.
  • Recognise recommendation, association-rule and time-series approaches and their assumptions.
  • Identify modelling limitations and prepare for further practice or separately administered certification assessment.
Prerequisites

Prerequisites

  • Basic computer, file-management and internet skills and high-school-level mathematics.
  • A supported Python 3 environment with compatible packages or a supplied notebook setup; approved installation rights only where needed.
  • For remote delivery, a reliable connection and conferencing equipment; a second screen is helpful, not compulsory.
  • A Google account is needed only if the chosen lab uses Colab; no unrestricted root access or fixed attendance-based certification promise is implied.
Training outline

7 modules

·
01Day 1 — Programming and data-science context3 topics
  • Set up Python and packages; variables, collections, loops, functions, scope and basic object-oriented concepts.
  • Use exceptions and simple text/regular-expression examples; consult documentation rather than memorise history.
  • Introduce R concepts through a brief comparison; the hands-on progression is primarily Python, not a full second-language course.
02Day 2 — Arrays, tables and statistics4 topics
  • NumPy arrays, matrices and selected operations; pandas DataFrames and CSV import/export.
  • Descriptive statistics, distributions and basic probability; linear-algebra concepts needed for the examples.
  • Introduce inferential statistics and hypothesis tests with assumptions and interpretation limits.
  • Practical work: inspect and summarise a small dataset.
03Day 3 — Preprocessing and visualisation4 topics
  • Missing values, categorical encoding, scaling and outlier decisions.
  • Separate training and evaluation data before fitting transformations; use an appropriate pipeline to reduce leakage.
  • Create selected Matplotlib or Seaborn charts and check whether they support the analytical question.
  • Build a documented preparation template and inspect its effect on the sample data.
04Day 4 — Regression and evaluation4 topics
  • Simple, multiple and polynomial regression: intuition, suitable problem types and assumptions.
  • Introduce decision-tree and random-forest regression and compare selected implementations.
  • Choose meaningful evaluation measures, compare a baseline and recognise overfitting.
  • Review results without implying causal explanation or guaranteed predictive performance.
05Day 5 — Classification4 topics
  • Compare logistic regression, trees, random forests, SVM, naïve Bayes and k-nearest neighbours at an introductory level.
  • Implement selected classifiers; broader algorithms use worked demonstrations.
  • Use a confusion matrix and appropriate metrics; consider class imbalance and data limitations.
  • Review validation and model settings without using held-out data to choose preprocessing.
06Day 6 — Unsupervised and recommendation methods4 topics
  • K-means and hierarchical clustering; inspect scaling, cluster assumptions and interpretation.
  • PCA for unsupervised dimensionality reduction; distinguish supervised linear discriminant analysis.
  • Content-based and collaborative recommendation approaches; assess a small supplied example.
  • Apriori and market-basket association rules; distinguish association from causation.
07Day 7 — Time series, consolidation and preparation4 topics
  • Time-series structure, basic ARIMA concepts and chronological evaluation using a small example.
  • Review statistical assumptions, model residuals and limitations.
  • Complete a bounded analysis exercise and document preparation, evaluation and interpretation.
  • Consolidate concepts with practice questions; consult current provider scope and exam policies before separately registering.

Certification context

The course preserves the source’s certification-preparation context but is not represented as the provider’s current self-paced product or an official credential award. Confirm current exam scope, access and retake conditions directly with Global Tech Council.

A programme built around your team.

Share your training goals and requirements.

Data Science Developer Foundations and Certification Preparation
FA-0720

Share your requirements for this programme.

Training enquiry

Data Science Developer Foundations and Certification Preparation