Exploring the hospital dataset
Exploratory data analysis is a preliminary step prior to data modeling in which you look at all of the characteristics of data in order t0 get a sense of data distribution, correlation, missing values, outliers, and any other factors that might impact future analyses. It is a very important step, and if performed diligently, will save you a lot of time later on.
For the following examples, we will read the NYC hospital discharges dataset (hospital inpatient discharges (SPARCS De-Identified): 2012, n.d.). This example uses the read.csv
function to input the delimited file, and then uses the View
function to graphically display the output. Then the str function is used to display the contents of the df dataframe that was just created, and then finally, the summary()
function displays all of the relevant statistics on all of the variables. These are all typical first steps to perform when looking at data for the first time:
df <-read.csv("C:/PracticalPredictiveAnalytics...