Search icon CANCEL
Subscription
0
Cart icon
Your Cart (0 item)
Close icon
You have no products in your basket yet
Save more on your purchases! discount-offer-chevron-icon
Savings automatically calculated. No voucher code required.
Arrow left icon
Explore Products
Best Sellers
New Releases
Books
Videos
Audiobooks
Learning Hub
Newsletter Hub
Free Learning
Arrow right icon
timer SALE ENDS IN
0 Days
:
00 Hours
:
00 Minutes
:
00 Seconds
Apache Spark 2.x Machine Learning Cookbook
Apache Spark 2.x Machine Learning Cookbook

Apache Spark 2.x Machine Learning Cookbook: Over 100 recipes to simplify machine learning model implementations with Spark

Arrow left icon
Profile Icon Amirghodsi Profile Icon Mohammed Guller Profile Icon Shuen Mei Profile Icon Rajendran Profile Icon Hall +1 more Show less
Arrow right icon
€8.99 €32.99
Full star icon Full star icon Full star icon Empty star icon Empty star icon 3 (2 Ratings)
eBook Sep 2017 666 pages 1st Edition
eBook
€8.99 €32.99
Paperback
€41.99
Subscription
Free Trial
Renews at €18.99p/m
Arrow left icon
Profile Icon Amirghodsi Profile Icon Mohammed Guller Profile Icon Shuen Mei Profile Icon Rajendran Profile Icon Hall +1 more Show less
Arrow right icon
€8.99 €32.99
Full star icon Full star icon Full star icon Empty star icon Empty star icon 3 (2 Ratings)
eBook Sep 2017 666 pages 1st Edition
eBook
€8.99 €32.99
Paperback
€41.99
Subscription
Free Trial
Renews at €18.99p/m
eBook
€8.99 €32.99
Paperback
€41.99
Subscription
Free Trial
Renews at €18.99p/m

What do you get with eBook?

Product feature icon Instant access to your Digital eBook purchase
Product feature icon Download this book in EPUB and PDF formats
Product feature icon Access this title in our online reader with advanced features
Product feature icon DRM FREE - Read whenever, wherever and however you want
OR
Modal Close icon
Payment Processing...
tick Completed

Billing Address

Table of content icon View table of contents Preview book icon Preview Book

Apache Spark 2.x Machine Learning Cookbook

Just Enough Linear Algebra for Machine Learning with Spark

In this chapter, we will cover the following recipes:

  • Package imports and initial setup for vectors and matrices
  • Creating DenseVector and setup with Spark 2.0
  • Creating SparseVector and setup with Spark 2.0
  • Creating DenseMatrix and setup with Spark 2.0
  • Using sparse local matrices with Spark 2.0
  • Performing vector arithmetic using Spark 2.0
  • Performing matrix arithmetic with Spark 2.0
  • Distributed matrices in Spark 2.0 ML library
  • Exploring RowMatrix in Spark 2.0
  • Exploring distributed IndexedRowMatrix in Spark 2.0
  • Exploring distributed CoordinateMatrix in Spark 2.0
  • Exploring distributed BlockMatrix in Spark 2.0

Introduction

Linear algebra is the cornerstone of machine learning (ML) and mathematical programming (MP). When dealing with Spark's machine library, one must understand that the Vector/Matrix structures provided by Scala (imported by default) are different from the Spark ML, MLlib Vector, Matrix facilities provided by Spark. The latter, powered by RDDs, is the desired data structure if you are going to use Spark (that is, parallelism) out of the box for large-scale matrix/vector computation (for example, SVD implementation alternatives with more numerical accuracy, desired in some cases for derivatives pricing and risk analytics). The Scala Vector/Matrix libraries provide a rich set of linear algebra operations such as dot product, additions, and so on, that still have their own place in an ML pipeline. In summary, the key difference between using Scala Breeze and Spark...

Package imports and initial setup for vectors and matrices

Before we can program in Spark or use vector and matrix artifacts, we need to first import the right packages and then set up SparkSession so we can gain access to the cluster handle.

In this short recipe, we highlight a comprehensive number of packages that can cover most of the linear algebra operations in Spark. The individual recipes that follow will include the exact subset required for the specific program.

How to do it...

  1. Start a new project in IntelliJ or in an IDE of your choice. Make sure that the necessary JAR files are included.
  2. Set up the package location where the program will reside:
package spark.ml.cookbook.chapter2
  1.  Import the necessary packages...

Creating DenseVector and setup with Spark 2.0

In this recipe, we explore DenseVectors using the Spark 2.0 machine library.

Spark provides two distinct types of vector facilities (dense and sparse) for storing and manipulating feature vectors that are going to be used in machine learning or optimization algorithms.

How to do it...

  1. In this section, we examine DenseVector examples that you would most likely use for implementing/augmenting existing machine learning programs. These examples also help to better understand Spark ML or MLlib source code and the underlying implementation (for example, Single Value Decomposition).
  2. Here we look at creating an ML vector feature (with independent variables) from arrays, which is a common...

Creating SparseVector and setup with Spark

In this recipe, we examine several types of SparseVector creation. As the length of the vector increases (millions) and the density remains low (few non-zero members), then sparse representation becomes more and more advantageous over the DenseVector.

How to do it...

  1. Start a new project in IntelliJ or in an IDE of your choice. Make sure that the necessary JAR files are included.
  2. Import the necessary packages for vector and matrix manipulation:
import org.apache.spark.sql.{SparkSession}
import org.apache.spark.mllib.linalg._
import breeze.linalg.{DenseVector => BreezeVector}
import Array._
import org.apache.spark.mllib.linalg.SparseVector
  1. Set up the Spark context and application...

Creating dense matrix and setup with Spark 2.0

In this recipe, we explore matrix creation examples that you most likely would need in your Scala programming and while reading the source code for many of the open source libraries for machine learning.

Spark provides two distinct types of local matrix facilities (dense and sparse) for storage and manipulation of data at a local level. For simplicity, one way to think of a matrix is to visualize it as columns of Vectors.

Getting ready

The key to remember here is that the recipe covers local matrices stored on one machine. We will use another recipe, Distributed matrices in the Spark2.0 ML library, covered in this chapter, for storing and manipulating distributed matrices.

...

Using sparse local matrices with Spark 2.0

In this recipe, we concentrate on SparseMatrix creation. In the previous recipe, we saw how a local dense matrix is declared and stored. A good number of machine learning problem domains can be represented as a set of features and labels within the matrix. In large-scale machine learning problems (for example, progression of a disease through large population centers, security fraud, political movement modeling, and so on), a good portion of the cells will be 0 or null (for example, the current number of people with a given disease versus the healthy population).

To help with storage and efficient operation in real time, sparse local matrices specialize in storing the cells efficiently as a list plus an index, which leads to faster loading and real time operations.

...

Introduction


Linear algebra is the cornerstone of machine learning (ML) and mathematicalprogramming (MP). When dealing with Spark's machine library, one must understand that the Vector/Matrix structures by Scala (imported by default) are different from the Spark ML, MLlib Vector, Matrix facilities provided by Spark. The latter, powered by RDDs, is the desired data structure if you are going to use Spark (that is, parallelism) out of the box for large-scale matrix/vector computation (for example, SVD implementation alternatives with more numerical accuracy, desired in some cases for derivatives pricing and risk analytics). The Scala Vector/Matrix libraries provide a rich set of linear algebra operations such as dot product, additions, and so on, that still have their own place in an ML pipeline. In summary, the key difference between using Scala Breeze and Spark or Spark ML is that the Spark facility is backed by RDDs which allows for simultaneous distributed, concurrent computing, and resiliency...

Package imports and initial setup for vectors and matrices


Before we can program in Spark or use and matrix artifacts, we need to first import the right packages and then set up SparkSession so we can gain access to the cluster handle.

In this short recipe, we highlight a comprehensive number of packages that can cover most of the linear algebra operations in Spark. The individual recipes that follow will include the subset required for the specific program.

How to do it...

  1. Start a new project in IntelliJ or in an IDE of your choice. Make sure that the necessary JAR files are included.
  2. Set up the package location where the program will reside:
package spark.ml.cookbook.chapter2
  1.  Import the necessary packages for vector and matrix manipulation:
import org.apache.spark.mllib.linalg.distributed.RowMatrix
import org.apache.spark.mllib.linalg.distributed.{IndexedRow, IndexedRowMatrix}
import org.apache.spark.mllib.linalg.distributed.{CoordinateMatrix, MatrixEntry}
import org.apache.spark.sql.{SparkSession...

Creating DenseVector and setup with Spark 2.0


In this recipe, we explore DenseVectors using the Spark 2.0 library.

Spark provides two types of vector facilities (dense and sparse) for storing and manipulating feature vectors that are going to be used in learning or optimization algorithms.

How to do it...

  1. In this section, we examine DenseVector examples that you would most likely use for implementing/augmenting existing machine learning programs. These examples also help to better understand Spark ML or MLlib source code and the underlying implementation (for example, Single Value Decomposition).
  2. Here we look at creating an ML vector feature (with independent variables) from arrays, which is a common use case. In this case, we have three almost fully populated Scala arrays corresponding to customer and product feature sets. We convert these arrays to the corresponding DenseVectors in Scala:
val CustomerFeatures1: Array[Double] = Array(1,3,5,7,9,1,3,2,4,5,6,1,2,5,3,7,4,3,4,1)
 val CustomerFeatures2...

Creating SparseVector and setup with Spark


In this recipe, we several types of SparseVector creation. As the length of the vector increases (millions) and the density remains low (few non-zero members), then sparse representation more and more advantageous over the DenseVector.

How to do it...

  1. Start a new project in IntelliJ or in an IDE of your choice. Make sure that the necessary JAR files are included.
  2. Import the necessary packages for vector and matrix manipulation:
import org.apache.spark.sql.{SparkSession}
import org.apache.spark.mllib.linalg._
import breeze.linalg.{DenseVector => BreezeVector}
import Array._
import org.apache.spark.mllib.linalg.SparseVector
  1. Set up the Spark context and application parameters so Spark can run. See the first recipe in this chapter for more details and variations:
val spark = SparkSession
 .builder
 .master("local[*]")
 .appName("myVectorMatrix")
 .config("spark.sql.warehouse.dir", ".")
 .getOrCreate()
  1. Here we look at creating a ML SparseVector that corresponds...

Creating dense matrix and setup with Spark 2.0


In this recipe, we explore creation examples that you most likely would need in your Scala programming and while reading the source code for many of the open source libraries for machine learning.

Spark provides two distinct types of local matrix facilities (dense and sparse) for storage and manipulation of data at a local level. For simplicity, one way to think of a is to visualize it as columns of Vectors.

Getting ready

The key to remember here is that the recipe covers local matrices stored on one machine. We will use another recipe, Distributed matrices in the Spark2.0 ML library, covered in this chapter, for storing and manipulating distributed matrices.

How to do it...

  1. Start a new project in IntelliJ or in an IDE of your choice. Make sure that the necessary JAR files are included.
  2. Import the necessary packages for vector and matrix manipulation:
 import org.apache.spark.sql.{SparkSession}
 import org.apache.spark.mllib.linalg._
 import breeze...

Using sparse local matrices with Spark 2.0


In this recipe, we concentrate on creation. In the recipe, we saw how a local dense matrix is declared and stored. A good number of machine learning problem domains can be represented as a set of features and labels within the matrix. In large-scale machine learning problems (for example, progression of a disease through large population centers, security fraud, political movement modeling, and so on), a good portion of the cells will be 0 or null (for example, the current number of people with a given disease versus the healthy population).

To help with storage and efficient operation in real time, sparse local matrices specialize in storing the cells efficiently as a list plus an index, which leads to faster loading and real time operations.

How to do it...

  1. Start a new project in IntelliJ or in an IDE of your choice. Make sure that the necessary JAR files are included.
  2. Import the necessary packages for vector and matrix manipulation: 
 import org...
Left arrow icon Right arrow icon
Download code icon Download Code

Key benefits

  • Solve the day-to-day problems of data science with Spark
  • This unique cookbook consists of exciting and intuitive numerical recipes
  • Optimize your work by acquiring, cleaning, analyzing, predicting, and visualizing your data

Description

Machine learning aims to extract knowledge from data, relying on fundamental concepts in computer science, statistics, probability, and optimization. Learning about algorithms enables a wide range of applications, from everyday tasks such as product recommendations and spam filtering to cutting edge applications such as self-driving cars and personalized medicine. You will gain hands-on experience of applying these principles using Apache Spark, a resilient cluster computing system well suited for large-scale machine learning tasks. This book begins with a quick overview of setting up the necessary IDEs to facilitate the execution of code examples that will be covered in various chapters. It also highlights some key issues developers face while working with machine learning algorithms on the Spark platform. We progress by uncovering the various Spark APIs and the implementation of ML algorithms with developing classification systems, recommendation engines, text analytics, clustering, and learning systems. Toward the final chapters, we’ll focus on building high-end applications and explain various unsupervised methodologies and challenges to tackle when implementing with big data ML systems.

Who is this book for?

This book is for Scala developers with a fairly good exposure to and understanding of machine learning techniques, but lack practical implementations with Spark. A solid knowledge of machine learning algorithms is assumed, as well as hands-on experience of implementing ML algorithms with Scala. However, you do not need to be acquainted with the Spark ML libraries and ecosystem.

What you will learn

  • Get to know how Scala and Spark go hand-in-hand for developers when developing ML systems with Spark
  • Build a recommendation engine that scales with Spark
  • Find out how to build unsupervised clustering systems to classify data in Spark
  • Build machine learning systems with the Decision Tree and Ensemble models in Spark
  • Deal with the curse of high-dimensionality in big data using Spark
  • Implement Text analytics for Search Engines in Spark
  • Streaming Machine Learning System implementation using Spark

Product Details

Country selected
Publication date, Length, Edition, Language, ISBN-13
Publication date : Sep 22, 2017
Length: 666 pages
Edition : 1st
Language : English
ISBN-13 : 9781782174608
Vendor :
Apache
Category :
Languages :

What do you get with eBook?

Product feature icon Instant access to your Digital eBook purchase
Product feature icon Download this book in EPUB and PDF formats
Product feature icon Access this title in our online reader with advanced features
Product feature icon DRM FREE - Read whenever, wherever and however you want
OR
Modal Close icon
Payment Processing...
tick Completed

Billing Address

Product Details

Publication date : Sep 22, 2017
Length: 666 pages
Edition : 1st
Language : English
ISBN-13 : 9781782174608
Vendor :
Apache
Category :
Languages :

Packt Subscriptions

See our plans and pricing
Modal Close icon
€18.99 billed monthly
Feature tick icon Unlimited access to Packt's library of 7,000+ practical books and videos
Feature tick icon Constantly refreshed with 50+ new titles a month
Feature tick icon Exclusive Early access to books as they're written
Feature tick icon Solve problems while you work with advanced search and reference features
Feature tick icon Offline reading on the mobile app
Feature tick icon Simple pricing, no contract
€189.99 billed annually
Feature tick icon Unlimited access to Packt's library of 7,000+ practical books and videos
Feature tick icon Constantly refreshed with 50+ new titles a month
Feature tick icon Exclusive Early access to books as they're written
Feature tick icon Solve problems while you work with advanced search and reference features
Feature tick icon Offline reading on the mobile app
Feature tick icon Choose a DRM-free eBook or Video every month to keep
Feature tick icon PLUS own as many other DRM-free eBooks or Videos as you like for just €5 each
Feature tick icon Exclusive print discounts
€264.99 billed in 18 months
Feature tick icon Unlimited access to Packt's library of 7,000+ practical books and videos
Feature tick icon Constantly refreshed with 50+ new titles a month
Feature tick icon Exclusive Early access to books as they're written
Feature tick icon Solve problems while you work with advanced search and reference features
Feature tick icon Offline reading on the mobile app
Feature tick icon Choose a DRM-free eBook or Video every month to keep
Feature tick icon PLUS own as many other DRM-free eBooks or Videos as you like for just €5 each
Feature tick icon Exclusive print discounts

Frequently bought together


Stars icon
Total 137.97
Scala and Spark for Big Data Analytics
€53.99
Mastering Machine Learning with Spark 2.x
€41.99
Apache Spark 2.x Machine Learning Cookbook
€41.99
Total 137.97 Stars icon
Banner background image

Table of Contents

13 Chapters
Practical Machine Learning with Spark Using Scala Chevron down icon Chevron up icon
Just Enough Linear Algebra for Machine Learning with Spark Chevron down icon Chevron up icon
Spark's Three Data Musketeers for Machine Learning - Perfect Together Chevron down icon Chevron up icon
Common Recipes for Implementing a Robust Machine Learning System Chevron down icon Chevron up icon
Practical Machine Learning with Regression and Classification in Spark 2.0 - Part I Chevron down icon Chevron up icon
Practical Machine Learning with Regression and Classification in Spark 2.0 - Part II Chevron down icon Chevron up icon
Recommendation Engine that Scales with Spark Chevron down icon Chevron up icon
Unsupervised Clustering with Apache Spark 2.0 Chevron down icon Chevron up icon
Optimization - Going Down the Hill with Gradient Descent Chevron down icon Chevron up icon
Building Machine Learning Systems with Decision Tree and Ensemble Models Chevron down icon Chevron up icon
Curse of High-Dimensionality in Big Data Chevron down icon Chevron up icon
Implementing Text Analytics with Spark 2.0 ML Library Chevron down icon Chevron up icon
Spark Streaming and Machine Learning Library Chevron down icon Chevron up icon

Customer reviews

Rating distribution
Full star icon Full star icon Full star icon Empty star icon Empty star icon 3
(2 Ratings)
5 star 50%
4 star 0%
3 star 0%
2 star 0%
1 star 50%
JIMMY WANG Nov 18, 2017
Full star icon Full star icon Full star icon Full star icon Full star icon 5
Great book! Very thorough with lots of examples. Will come back with more thoughts after finishing the whole book.
Amazon Verified review Amazon
taiwo Raphael Alabi Dec 06, 2018
Full star icon Empty star icon Empty star icon Empty star icon Empty star icon 1
Terrible book! DO not buy!
Amazon Verified review Amazon
Get free access to Packt library with over 7500+ books and video courses for 7 days!
Start Free Trial

FAQs

How do I buy and download an eBook? Chevron down icon Chevron up icon

Where there is an eBook version of a title available, you can buy it from the book details for that title. Add either the standalone eBook or the eBook and print book bundle to your shopping cart. Your eBook will show in your cart as a product on its own. After completing checkout and payment in the normal way, you will receive your receipt on the screen containing a link to a personalised PDF download file. This link will remain active for 30 days. You can download backup copies of the file by logging in to your account at any time.

If you already have Adobe reader installed, then clicking on the link will download and open the PDF file directly. If you don't, then save the PDF file on your machine and download the Reader to view it.

Please Note: Packt eBooks are non-returnable and non-refundable.

Packt eBook and Licensing When you buy an eBook from Packt Publishing, completing your purchase means you accept the terms of our licence agreement. Please read the full text of the agreement. In it we have tried to balance the need for the ebook to be usable for you the reader with our needs to protect the rights of us as Publishers and of our authors. In summary, the agreement says:

  • You may make copies of your eBook for your own use onto any machine
  • You may not pass copies of the eBook on to anyone else
How can I make a purchase on your website? Chevron down icon Chevron up icon

If you want to purchase a video course, eBook or Bundle (Print+eBook) please follow below steps:

  1. Register on our website using your email address and the password.
  2. Search for the title by name or ISBN using the search option.
  3. Select the title you want to purchase.
  4. Choose the format you wish to purchase the title in; if you order the Print Book, you get a free eBook copy of the same title. 
  5. Proceed with the checkout process (payment to be made using Credit Card, Debit Cart, or PayPal)
Where can I access support around an eBook? Chevron down icon Chevron up icon
  • If you experience a problem with using or installing Adobe Reader, the contact Adobe directly.
  • To view the errata for the book, see www.packtpub.com/support and view the pages for the title you have.
  • To view your account details or to download a new copy of the book go to www.packtpub.com/account
  • To contact us directly if a problem is not resolved, use www.packtpub.com/contact-us
What eBook formats do Packt support? Chevron down icon Chevron up icon

Our eBooks are currently available in a variety of formats such as PDF and ePubs. In the future, this may well change with trends and development in technology, but please note that our PDFs are not Adobe eBook Reader format, which has greater restrictions on security.

You will need to use Adobe Reader v9 or later in order to read Packt's PDF eBooks.

What are the benefits of eBooks? Chevron down icon Chevron up icon
  • You can get the information you need immediately
  • You can easily take them with you on a laptop
  • You can download them an unlimited number of times
  • You can print them out
  • They are copy-paste enabled
  • They are searchable
  • There is no password protection
  • They are lower price than print
  • They save resources and space
What is an eBook? Chevron down icon Chevron up icon

Packt eBooks are a complete electronic version of the print edition, available in PDF and ePub formats. Every piece of content down to the page numbering is the same. Because we save the costs of printing and shipping the book to you, we are able to offer eBooks at a lower cost than print editions.

When you have purchased an eBook, simply login to your account and click on the link in Your Download Area. We recommend you saving the file to your hard drive before opening it.

For optimal viewing of our eBooks, we recommend you download and install the free Adobe Reader version 9.