Subscription

Explore Products

Best Sellers

New Releases

Books

Videos

Audiobooks

Learning Hub

Conferences

Free Learning

You're reading from Hands-On Data Science with R Techniques to perform data manipulation and mining to build smart analytical models using R

Product type Paperback

Published in Nov 2018

Publisher Packt

ISBN-13 9781789139402

Length 420 pages

Edition 1st Edition

Languages

Tools

ggplot

Concepts

Data Science

Authors (4):

Nataraj Dasgupta

Vitor Bianchi Lanzetta

Doug Ortiz

Ricardo Anjoleto Farias

View More author details

Table of Contents (16) Chapters

Preface

1. Getting Started with Data Science and R FREE CHAPTER

2. Descriptive and Inferential Statistics

3. Data Wrangling with R

4. KDD, Data Mining, and Text Mining

5. Data Analysis with R

6. Machine Learning with R

7. Forecasting and ML App with R

8. Neural Networks and Deep Learning

9. Markovian in R

10. Visualizing Data

11. Going to Production with R

12. Large Scale Data Analytics with Hadoop

13. R on Cloud

14. The Road Ahead

15. Other Books You May Enjoy

Leave a review - let other readers know what you think

Manipulating Spark data using both dplyr and SQL

Once you're done with the installation from this chapters introduction, let's create a remote dplyr data source for the Spark cluster. To do this, use the spark_connect function, as shown:

sc <- spark_connect(master = "local")

This will create a Spark cluster in your computer; you can see it at your RStudio, a tab guide alongside your R environment. To disconnect, use the spark_disconnect(sc) function. Keep connected and copy a couple of datasets from any R packages into the cluster:

library(DAAG)
dt_sugar <- copy_to(sc, sugar, "SUGAR")
dt_stVincent <- copy_to(sc, stVincent, "STVINCENT")

The preceding code uploads the DAAG::sugar and DAAG::stVicent DataFrames into the your connected Spark cluster. It also creates the table definitions; they were saved into dt_sugar and dt_stVincent...

The rest of the chapter is locked

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at €18.99/month. Cancel anytime

Authors (4)

Dasgupta

Nataraj Dasgupta is the vice president of advanced analytics at RxDataScience Inc. Nataraj has been in the IT industry for more than 19 years, and has worked in the technical and analytics divisions of Philip Morris, IBM, UBS Investment Bank, and Purdue Pharma. At Purdue Pharma, Nataraj led the data science division, where he developed the company's award-winning big data and machine learning platform. Prior to Purdue, at UBS, he held the role of Associate Director, working with high-frequency and algorithmic trading technologies in the foreign exchange trading division of the bank.

See other products by Dasgupta

Bianchi Lanzetta

Vitor Bianchi Lanzetta (@vitorlanzetta) has a master's degree in Applied Economics (University of So PauloUSP) and works as a data scientist in a tech start-up named RedFox Digital Solutions. He has also authored a book called R Data Visualization Recipes. The things he enjoys the most are statistics, economics, and sports of all kinds (electronics included). His blog, made in partnership with Ricardo Anjoleto Farias (@R_A_Farias), can be found at ArcadeData dot org, they kindly call it R-Cade Data.

See other products by Bianchi Lanzetta

Doug Ortiz

Doug Ortiz is an experienced enterprise cloud, big data, data analytics, and solutions architect who has architected, designed, developed, engineered, re-engineered, and integrated enterprise solutions. The technologies he has experience with include: Amazon Web Services, Azure, Google Cloud, Business Intelligence, Data Science, Hadoop, Spark, NoSQL and Graph Databases, and Web Front-End Technologies.

See other products by Doug Ortiz

Farias

Ricardo Anjoleto Farias is an economist who graduated from the Universidade Estadual de Maring in 2014. In addition to being a sports enthusiast (electronic or otherwise) and enjoying a good barbecue, he also likes math, statistics, and correlated studies. His first contact with R was when he embarked on his master's degree, and since then, he has tried to improve his skills with this powerful tool.

See other products by Farias