Subscription

Explore Products

Best Sellers

New Releases

Books

Videos

Audiobooks

Learning Hub

Newsletter Hub

Free Learning

You're reading from Mastering Machine Learning with Spark 2.x Harness the potential of machine learning, through spark

Product type Paperback

Published in Aug 2017

Publisher Packt

ISBN-13 9781785283451

Length 340 pages

Edition 1st Edition

Languages

Java

Tools

Apache Spark

Concepts

Machine Learning

Authors (3):

Malohlava

Tellez

Max Pumperla

View More author details

Table of Contents (9) Chapters

Preface

1. Introduction to Large-Scale Machine Learning and Spark FREE CHAPTER

2. Detecting Dark Matter - The Higgs-Boson Particle

3. Ensemble Methods for Multi-Class Classification

4. Predicting Movie Reviews Using NLP and Spark Streaming

5. Word2vec for Prediction and Clustering

6. Extracting Patterns from Clickstream Data

7. Graph Analytics with GraphX

8. Lending Club Loan Prediction

Featurization - feature hashing

Now, it is time to transform string representation into a numeric one. We adopt a bag-of-words approach; however, we use a trick called feature hashing. Let's look in more detail at how Spark employs this powerful technique to help us construct and access our tokenized dataset efficiently. We use feature hashing as a time-efficient implementation of a bag-of-words, as explained earlier.

At its core, feature hashing is a fast and space-efficient method to deal with high-dimensional data-typical in working with text-by converting arbitrary features into indices within a vector or matrix. This is best described with an example text. Suppose we have the following two movie reviews:

The movie Goodfellas was well worth the money spent. Brilliant acting!
Goodfellas is a riveting movie with a great cast and a brilliant plot-a must see for all...

The rest of the chapter is locked

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at €18.99/month. Cancel anytime

Authors (3)

Malohlava

Michal Malohlava, creator of Sparkling Water, is a geek and the developer; Java, Linux, programming languages enthusiast who has been developing software for over 10 years. He obtained his PhD from Charles University in Prague in 2012, and post doctorate from Purdue University. During his studies, he was interested in the construction of not only distributed but also embedded and real-time, component-based systems, using model-driven methods and domain-specific languages. He participated in the design and development of various systems, including SOFA and Fractal component systems and the jPapabench control system. Now, his main interest is big data computation. He participates in the development of the H2O platform for advanced big data math and computation, and its embedding into Spark engine, published as a project called Sparkling Water.

See other products by Malohlava

Tellez

Alex Tellez is a life-long data hacker/enthusiast with a passion for data science and its application to business problems. He has a wealth of experience working across multiple industries, including banking, health care, online dating, human resources, and online gaming. Alex has also given multiple talks at various AI/machine learning conferences, in addition to lectures at universities about neural networks. When hes not neck-deep in a textbook, Alex enjoys spending time with family, riding bikes, and utilizing machine learning to feed his French wine curiosity!

See other products by Tellez

Max Pumperla

Max Pumperla is a data scientist and engineer specializing in deep learning and its applications. He currently works as a deep learning engineer at Skymind and is a co-founder of aetros.com. Max is the author and maintainer of several Python packages, including elephas, a distributed deep learning library using Spark. His open source footprint includes contributions to many popular machine learning libraries, such as keras, deeplearning4j, and hyperopt. He holds a PhD in algebraic geometry from the University of Hamburg.

See other products by Max Pumperla