Packt+ | Advance your knowledge in tech

You're reading from Python Machine Learning Learn how to build powerful Python machine learning algorithms to generate useful data insights with this data analysis tutorial

Product type Paperback

Published in Sep 2015

Publisher Packt

ISBN-13 9781783555130

Length 454 pages

Edition 1st Edition

Languages

Python

Tools

SciPy

Concepts

Machine Learning

Author (1):

Sebastian Raschka

View More author details

Table of Contents (15) Chapters

Preface

1. Giving Computers the Ability to Learn from Data FREE CHAPTER

2. Training Machine Learning Algorithms for Classification

3. A Tour of Machine Learning Classifiers Using Scikit-learn

4. Building Good Training Sets – Data Preprocessing

5. Compressing Data via Dimensionality Reduction

6. Learning Best Practices for Model Evaluation and Hyperparameter Tuning

7. Combining Different Models for Ensemble Learning

8. Applying Machine Learning to Sentiment Analysis

9. Embedding a Machine Learning Model into a Web Application

10. Predicting Continuous Target Variables with Regression Analysis

11. Working with Unlabeled Data – Clustering Analysis

12. Training Artificial Neural Networks for Image Recognition

13. Parallelizing Neural Network Training with Theano

Index

Evaluating the performance of linear regression models

In the previous section, we discussed how to fit a regression model on training data. However, you learned in previous chapters that it is crucial to test the model on data that it hasn't seen during training to obtain an unbiased estimate of its performance.

As we remember from Chapter 6, Learning Best Practices for Model Evaluation and Hyperparameter Tuning, we want to split our dataset into separate training and test datasets where we use the former to fit the model and the latter to evaluate its performance to generalize to unseen data. Instead of proceeding with the simple regression model, we will now use all variables in the dataset and train a multiple regression model:

>>> from sklearn.cross_validation import train_test_split
>>> X = df.iloc[:, :-1].values
>>> y = df['MEDV'].values
>>> X_train, X_test, y_train, y_test = train_test_split(
...       X, y, test_size=0.3, random_state=0)
>&gt...