Business case
In the upcoming case study, we will apply KNN and SVM to the same dataset. This will allow us to compare the R code and learning methods on the same problem, starting with KNN. We will also spend some time drilling down into the confusion matrix, comparing a number of statistics to evaluate model accuracy.
Business understanding
The data that we will examine was originally collected by the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK). It consists of 532
observations and eight input features along with a binary outcome (Yes
/No
). The patients in this study were of Pima Indian descent from South Central Arizona. The NIDDK data shows that for the last 30 years, research has helped scientists to prove that obesity is a major risk factor in the development of diabetes. The Pima Indians were selected for the study as one-half of the adult Pima Indians have diabetes and 95 per cent of those with diabetes are overweight. The analysis will focus on adult women...