You're reading from Machine Learning with R Learn techniques for building and improving machine learning models, from data preparation to model tuning, evaluation, and working with big data

Product type Paperback

Published in May 2023

Publisher Packt

ISBN-13 9781801071321

Length 762 pages

Edition 4th Edition

Languages

Tools

H2O

Concepts

Big Data

Author (1):

Brett Lantz

View More author details

Table of Contents (18) Chapters

Preface

1. Introducing Machine Learning

2. Managing and Understanding Data FREE CHAPTER

3. Lazy Learning – Classification Using Nearest Neighbors

4. Probabilistic Learning – Classification Using Naive Bayes

5. Divide and Conquer – Classification Using Decision Trees and Rules

6. Forecasting Numeric Data – Regression Methods

7. Black-Box Methods – Neural Networks and Support Vector Machines

8. Finding Patterns – Market Basket Analysis Using Association Rules

9. Finding Groups of Data – Clustering with k-means

10. Evaluating Model Performance

11. Being Successful with Machine Learning

12. Advanced Data Preparation

13. Challenging Data – Too Much, Too Little, Too Complex

14. Building Better Learners

15. Making Use of Big Data

16. Other Books You May Enjoy

17. Index

Making use of sparse data

As datasets increase in dimension, some attributes are likely to be sparse, which means most observations do not share values of the attribute. This is a natural consequence of the curse of dimensionality in which this ever-increasing detail turns observations into outliers identified by their unique combination of attributes. It is very uncommon for sparse data to have any specific value, or perhaps even any value at all—as was the case in the sparse matrices for text data found in Chapter 4, Probabilistic Learning – Classification Using Naive Bayes, and the sparse matrices for shopping cart data in Chapter 8, Finding Patterns – Market Basket Analysis Using Association Rules.

This is not the same as missing data, where typically a relatively small portion of values are unknown. In sparse data, most values are known, but the number of interesting, meaningful values is dwarfed by an overwhelming number of values that add little value...