This blog post is related to Chapter 4 of "Practical Machine Learning with R", but it can also be read as a standalone article.
The KNIME platform is a freely available software. The KNIME workflow below repeats the k-Nearest Neighbors wine color analysis from Chapter 4 of Practical Machine Learning with R (https:\\ai.lange-analytics.com). We use the same training and testing data as in Chapter 4.
If you have KNIME installed you can download the workflow to study it and to make modifications.
Below is a diagram of the KNIME workflow. You can load the workflow to a local KNIME installation from this link: here
In the first two nodes on the left of the diagram, we load the training and testing data for the analysis. Training and testing data are loaded separately rather than loading one dataset and splitting it into training and testing data to ensure that we work with exactly the same data as in Chapter 4.
Then we use z-Score Normalization for the training and testing data as well. This is needed to give variables with different magnitudes a similar impact on the distance to the predicted observation.
Both training data and testing data are inputted into the k-Nearest Neighbor Learner, which finds the four nearest Neighbors for each of the testing observations from the training observations. The prediction is the majority of the wine colors of these four nearest Neighbors (in the case of a tie the wine color is picked randomly).
In the last node, the scorer generates a confusion matrix and calculates metrics such as accuracy (99.0%), sensitivity (99.2%), and specificity (98.8%).
The results are very similar to the one from Chapter 4.9 of "Practical Machine Learning with R".
Happy Analytics

Comments
Post a Comment