You're reading from Python Machine Learning Learn how to build powerful Python machine learning algorithms to generate useful data insights with this data analysis tutorial

Product type Paperback

Published in Sep 2015

Publisher Packt

ISBN-13 9781783555130

Length 454 pages

Edition 1st Edition

Languages

Python

Tools

SciPy

Concepts

Machine Learning

Author (1):

Sebastian Raschka

View More author details

Table of Contents (15) Chapters

Preface

1. Giving Computers the Ability to Learn from Data

2. Training Machine Learning Algorithms for Classification FREE CHAPTER

3. A Tour of Machine Learning Classifiers Using Scikit-learn

4. Building Good Training Sets – Data Preprocessing

5. Compressing Data via Dimensionality Reduction

6. Learning Best Practices for Model Evaluation and Hyperparameter Tuning

7. Combining Different Models for Ensemble Learning

8. Applying Machine Learning to Sentiment Analysis

9. Embedding a Machine Learning Model into a Web Application

10. Predicting Continuous Target Variables with Regression Analysis

11. Working with Unlabeled Data – Clustering Analysis

12. Training Artificial Neural Networks for Image Recognition

13. Parallelizing Neural Network Training with Theano

Index

Chapter 5. Compressing Data via Dimensionality Reduction

In Chapter 4, Building Good Training Sets – Data Preprocessing, you learned about the different approaches for reducing the dimensionality of a dataset using different feature selection techniques. An alternative approach to feature selection for dimensionality reduction is feature extraction. In this chapter, you will learn about three fundamental techniques that will help us to summarize the information content of a dataset by transforming it onto a new feature subspace of lower dimensionality than the original one. Data compression is an important topic in machine learning, and it helps us to store and analyze the increasing amounts of data that are produced and collected in the modern age of technology. In this chapter, we will cover the following topics:

Principal component analysis (PCA) for unsupervised data compression
Linear Discriminant Analysis (LDA) as a supervised dimensionality reduction technique for...