FA-0778Data & AnalyticsSoftware Development

Database, Data Analytics and Machine Learning Foundations

Python data acquisition, SQL Server and guided analysis

Introduction

Why this course

Build a foundation in relational data, Python-based analysis and introductory machine learning. Work with prepared files, APIs and database connections, then use guided examples to clean, visualise and model data. The course combines basic programming and database skills with selected machine-learning demonstrations; it is not advanced mastery of arbitrary data sources or production infrastructure.

Learning outcomes

Learning outcomes

  • Describe SDLC phases and development practices in a data project.
  • Use basic Python, external libraries, files and encoding for data acquisition and preparation.
  • Connect to a prepared relational database and run basic SQL queries.
  • Visualise data and apply selected classification, clustering and regression methods.
  • Evaluate models using separated training/test data and cross-validation.
  • Recognise Linux, SSH and remote-query concepts used in the prepared environment.
Prerequisites

Prerequisites

Basic computer, internet and file-handling skills and English-language technical reading. Use a prepared compatible environment and authorised access to course data and remote services; only permissions needed for the exercises are required.

Training outline

4 modules

·
01Day 1 — Python, project lifecycle and data acquisition6 topics
  • SDLC phases, pre-development decisions and project-independent development practices.
  • Computing/workflow design concepts illustrated through a prepared data process.
  • Python syntax and basic data structures using a supported Python3 version.
  • External libraries, files, encoding, extraction and conversion.
  • Data acquisition through prepared files, permitted scraping/API examples and database connectivity.
  • Linux, SSH and authorised remote-query basics.
02Day 2 — SQL Server environment and SQL7 topics
  • RDBMS and Microsoft SQL Server; SQL and T-SQL.
  • Container and virtual-machine concepts; prepared supported Linux container rather than Ubuntu18.04 installation instructions.
  • Remote access, environment security and differences between lab and production setup.
  • Use SSMS or a supported client such as VS Code MSSQL; Azure Data Studio is retired.
  • Databases, tables, stored procedures, functions and import/export.
  • SELECT, WHERE, LIKE, ordering, INSERT, UPDATE, DELETE, IN and BETWEEN.
  • Aggregates, grouping, ALTER and subqueries.
03Day 3 — Preparation, visualisation and probability4 topics
  • Data visualisation with Seaborn
  • Covariance and Correlation.
  • Conditional Probability.
  • Bayes' Theorem.
  • Data Cleaning and Normalization
  • Cleaning web log data
  • Normalizing numerical data
  • Detecting outliers
  • Feature Engineering
  • Imputation Techniques for Missing Data
  • Handling Unbalanced Data: Oversampling

Fit imputation, scaling and other learned preparation only on training data; oversampling belongs within the training process, not the held-out evaluation data.

04Day 4 — Introductory machine learning11 topics
  • Basic Machine Learning (using Python).
  • Supervised vs. Unsupervised Learning, and Train/Test.
  • Use train/test evaluation to assess polynomial-regression overfitting
  • Bayesian Methods: Concepts.
  • Practice with:
  • Implementing a Spam Classifier with Naive Bayes.
  • K-Means Clustering.
  • Entropy
  • Decision Trees
  • Bias/Variance Tradeoff
  • K-fold cross-validation to assess generalisation and guide model selection

Introduce the roles of pandas, SciPy, scikit-learn, Matplotlib, Seaborn and XGBoost in the examples. Review predictions and complete the assessment.

A programme built around your team.

Share your training goals and requirements.

Database, Data Analytics and Machine Learning Foundations
FA-0778

Share your requirements for this programme.

Training enquiry

Database, Data Analytics and Machine Learning Foundations