Packt+ | Advance your knowledge in tech

You're reading from Python Machine Learning, Second Edition Machine Learning and Deep Learning with Python, scikit-learn, and TensorFlow

Product type Paperback

Published in Sep 2017

Publisher Packt

ISBN-13 9781787125933

Length 622 pages

Edition 2nd Edition

Languages

Python

Tools

Scikit-learn

Concepts

Deep Learning

Authors (2):

Vahid Mirjalili

Sebastian Raschka

View More author details

Table of Contents (18) Chapters

Preface

1. Giving Computers the Ability to Learn from Data FREE CHAPTER

2. Training Simple Machine Learning Algorithms for Classification

3. A Tour of Machine Learning Classifiers Using scikit-learn

4. Building Good Training Sets – Data Preprocessing

5. Compressing Data via Dimensionality Reduction

6. Learning Best Practices for Model Evaluation and Hyperparameter Tuning

7. Combining Different Models for Ensemble Learning

8. Applying Machine Learning to Sentiment Analysis

9. Embedding a Machine Learning Model into a Web Application

10. Predicting Continuous Target Variables with Regression Analysis

11. Working with Unlabeled Data – Clustering Analysis

12. Implementing a Multilayer Artificial Neural Network from Scratch

13. Parallelizing Neural Network Training with TensorFlow

14. Going Deeper – The Mechanics of TensorFlow

15. Classifying Images with Deep Convolutional Neural Networks

16. Modeling Sequential Data Using Recurrent Neural Networks

Index

Dealing with class imbalance

We've mentioned class imbalances several times throughout this chapter, and yet we haven't actually discussed how to deal with such scenarios appropriately if they occur. Class imbalance is a quite common problem when working with real-world data—samples from one class or multiple classes are over-represented in a dataset. Intuitively, we can think of several domains where this may occur, such as spam filtering, fraud detection, or screening for diseases.

Imagine the breast cancer dataset that we've been working with in this chapter consisted of 90 percent healthy patients. In this case, we could achieve 90 percent accuracy on the test dataset by just predicting the majority class (benign tumor) for all samples, without the help of a supervised machine learning algorithm. Thus, training a model on such a dataset that achieves approximately 90 percent test accuracy would mean our model hasn't learned anything useful from the features provided in this dataset.

In...