Subscription

Explore Products

Best Sellers

New Releases

Books

Events

Videos

Audiobooks

Packt Hub

Free Learning

You're reading from Learning Data Mining with Python Use Python to manipulate data and build predictive models

Product type Paperback

Published in Apr 2017

Publisher Packt

ISBN-13 9781787126787

Length 358 pages

Edition 2nd Edition

Languages

Python

Concepts

Data Mining

Author (1):

Robert Layton

View More author details

Table of Contents (14) Chapters

Preface

1. Getting Started with Data Mining

2. Classifying with scikit-learn Estimators FREE CHAPTER

3. Predicting Sports Winners with Decision Trees

4. Recommending Movies Using Affinity Analysis

5. Features and scikit-learn Transformers

6. Social Media Insight using Naive Bayes

7. Follow Recommendations Using Graph Mining

8. Beating CAPTCHAs with Neural Networks

9. Authorship Attribution

10. Clustering News Articles

11. Object Detection in Images using Deep Neural Networks

12. Working with Big Data

13. Next Steps...

Summary

In this chapter, we used several of scikit-learn's methods for building a standard workflow to run and evaluate data mining models. We introduced the Nearest Neighbors algorithm, which is implemented in scikit-learn as an estimator. Using this class is quite easy; first, we call the fit function on our training data, and second, we use the predict function to predict the class of testing samples.

We then looked at pre-processing by fixing poor feature scaling. This was done using a Transformer object and the MinMaxScaler class. These functions also have a fit method and then a transform, which takes data of one form as an input and returns a transformed dataset as an output.

To investigate these transformations further, try swapping out the MinMaxScaler with some of the other mentioned transformers. Which is the most effective and why would this be the case?

Other transformers also exist in scikit-learn...

Tech Concepts

Programming languages

Tech Tools

Unlimited access to the largest independent learning library in tech of over 8,000 expert-authored tech books and videos.

Innovative learning tools, including AI book assistants, code context explainers, and text-to-speech.

50+ new titles added per month and exclusive early access to books as they are being written.

You have been reading a chapter from

Learning Data Mining with Python - Second Edition

Published in: Apr 2017

Publisher: Packt

ISBN-13: 9781787126787

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at $19.99/month. Cancel anytime

Authors (1)

Robert Layton

Robert Layton is a data scientist investigating data-driven applications to businesses across a number of sectors. He received a PhD investigating cybercrime analytics from the Internet Commerce Security Laboratory at Federation University Australia, before moving into industry, starting his own data analytics company dataPipeline. Next, he created Eureaktive, which works with tech-based startups on developing their proof-of-concepts and early-stage prototypes. Robert also runs the LearningTensorFlow website, which is one of the world's premier tutorial websites for Google's TensorFlow library. Robert is an active member of the Python community, having used Python for more than 8 years. He has presented at PyConAU for the last four years and works with Python Charmers to provide Python-based training for businesses and professionals from a wide range of organisations. Robert can be best reached via Twitter @robertlayton

See other products by Robert Layton