Course Outline for Data Science with KNIME
The goal of this course is to enable existing KNIE users to take their Data Analytics skills to the next level. This course starts with the core understanding of how Machine Learning works and how KNIME can be used to implement its algorithms on data.
This is not a beginners course and is meant for learners with experience using KNIME.
Course Outcome
By the end of this course, the learner may be expected to know and understand the following:
- Data Preparation
- Data Cleaning
- Data Preprocessing
- Modeling of algorithms
- Classification of data
- Machine Learning and its techniques
- ML methodologies for implementation
- Mining
- Prediction of targets
- Evaluation
Please Note: The outcome listed above is a guideline and not a guarantee. Every learner is unique and so is their learning ability. Variations in the ability to comprehend and utilize shall differ from pupil to pupil.
Prerequisites
As mentioned previously, this course is not intended for learners who are not KNIME users. In order to be able to participate in this course, the learner needs to fulfil the following prerequisites:
- Access to PC [either one of these OS: MS Windows / Mac / Linux (ubuntu)]
- Internet connection
- Webcam
- Microphone
- Dual screens
- Ability to comprehend the English language
- Minimum 99% attendance
- Basic education in mathematics (understanding of linear and non-linear mathematics at high-school level)
- Very basic understanding of API and web requests
- Experience and understanding of filing system
- Experience in Using KNIME for analytics
Course Outline
This two day course will cover the following topics:
- Module One
- Understanding of KNIME’s workflow
- Overview of workflow automation
- Basic flat file data manipulation
- Basic API connectivity and conjunction
- Module Two
- Recursion
- Implementation of Recursion on historic data
- Historic data analysis
- Basis of Machine Learning
- Module Three
- Simple Machine Learning Implementation
- Prediction
- Parameter Optimization
- Feature Optimization
- Module Four
- Linear Regression
- Evaluation of Regression Models
- Polynomial Regression
- Simple Regression Tree
- Module Five
- Bayes Theorem and Naive Bayes Model
- Decision Tree
- k-Nearest Neighbor Algorithm
- Distance Metrics
- SVM: Support Vector Machines
- Logistic Regression
- Comparing Classification Algorithms
- Module Six
- Introduction to Clustering and Concepts
- K-Means
- Optimizing number of clusters (k value) in Knime with Grid Search
- Hierarchical Clustering: Agglomerative and Divisive Approaches
- DBSCAN : Density Based Approach
- Module Seven
- PNML
- Assessment
Practical, connected learning
My wider training approach brings hands-on implementation and systems thinking together, connecting technology with real operational needs.