Summary
In this chapter, we extended our machine learning on Spark to serve learning analytics, for which we completed a step-by-step process of processing big data obtained from learning management systems and other sources for a rapid development of student attrition prediction models on Apache Spark. With the machine learning results obtained, we developed rules and scores to be used by NIY University for interventions to reduce student attrition.
Specifically, we first selected a supervised machine learning approach with a focus on logistic regression and decision trees as per the special needs of this university and the nature of the project, and after this, we prepared Spark computing and loaded in the preprocessed data. Secondly, we worked on feature development and selection. Thirdly, we estimated model coefficients with the Zeppeline notebook on Spark. Next, we evaluated these estimated models using a confusion matrix and error ratios. Then, we interpreted our machine learning results...