Explore Products

Best Sellers

New Releases

Books

Videos

Audiobooks

Learning Hub

Conferences

Free Learning

You're reading from Hands-On Data Science with R Techniques to perform data manipulation and mining to build smart analytical models using R

Product type Paperback

Published in Nov 2018

Publisher Packt

ISBN-13 9781789139402

Length 420 pages

Edition 1st Edition

Languages

Tools

ggplot

Concepts

Data Science

Authors (4):

Nataraj Dasgupta

Vitor Bianchi Lanzetta

Doug Ortiz

Ricardo Anjoleto Farias

View More author details

Table of Contents (16) Chapters

Preface

1. Getting Started with Data Science and R FREE CHAPTER

2. Descriptive and Inferential Statistics

3. Data Wrangling with R

4. KDD, Data Mining, and Text Mining

5. Data Analysis with R

6. Machine Learning with R

7. Forecasting and ML App with R

8. Neural Networks and Deep Learning

9. Markovian in R

10. Visualizing Data

11. Going to Production with R

12. Large Scale Data Analytics with Hadoop

13. R on Cloud

14. The Road Ahead

15. Other Books You May Enjoy

Leave a review - let other readers know what you think

Cleaning and transforming data

In Chapter 3, Data Wrangling with R, we approached the topic of data cleaning (munging). Data cleaning is so important that the majority of data scientists spend most of their work time cleaning and preparing data. The last session, What is the R community tweeting about?, gave us a DataFrame with 15999 rows and 42 columns. That is raw data. This session will clean and transform it.

Our initial goal was to check which packages the R community is talking about on Twitter. There are three variables we will use to achieve the final goal.

The variable text can be truncated when there is a retweet. When that is the case, check retweet_text, which won't be truncated. The quoted_text variable also brings useful information. To unite all the useful information into a single object, we can use the following code:

quotes <- tweets_dt$is_quote
rts...

The rest of the chapter is locked

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at $19.99/month. Cancel anytime

Authors (4)

Dasgupta

Nataraj Dasgupta is the vice president of advanced analytics at RxDataScience Inc. Nataraj has been in the IT industry for more than 19 years, and has worked in the technical and analytics divisions of Philip Morris, IBM, UBS Investment Bank, and Purdue Pharma. At Purdue Pharma, Nataraj led the data science division, where he developed the company's award-winning big data and machine learning platform. Prior to Purdue, at UBS, he held the role of Associate Director, working with high-frequency and algorithmic trading technologies in the foreign exchange trading division of the bank.

See other products by Dasgupta

Bianchi Lanzetta

Vitor Bianchi Lanzetta (@vitorlanzetta) has a master's degree in Applied Economics (University of So PauloUSP) and works as a data scientist in a tech start-up named RedFox Digital Solutions. He has also authored a book called R Data Visualization Recipes. The things he enjoys the most are statistics, economics, and sports of all kinds (electronics included). His blog, made in partnership with Ricardo Anjoleto Farias (@R_A_Farias), can be found at ArcadeData dot org, they kindly call it R-Cade Data.

See other products by Bianchi Lanzetta

Doug Ortiz

Doug Ortiz is an experienced enterprise cloud, big data, data analytics, and solutions architect who has architected, designed, developed, engineered, re-engineered, and integrated enterprise solutions. The technologies he has experience with include: Amazon Web Services, Azure, Google Cloud, Business Intelligence, Data Science, Hadoop, Spark, NoSQL and Graph Databases, and Web Front-End Technologies.

See other products by Doug Ortiz

Farias

Ricardo Anjoleto Farias is an economist who graduated from the Universidade Estadual de Maring in 2014. In addition to being a sports enthusiast (electronic or otherwise) and enjoying a good barbecue, he also likes math, statistics, and correlated studies. His first contact with R was when he embarked on his master's degree, and since then, he has tried to improve his skills with this powerful tool.

See other products by Farias