Packt+ | Advance your knowledge in tech

You're reading from Big Data Analysis with Python Combine Spark and Python to unlock the powers of parallel computing and machine learning

Product type Paperback

Published in Apr 2019

Publisher Packt

ISBN-13 9781789955286

Length 276 pages

Edition 1st Edition

Languages

Python

Tools

Combine

Concepts

Big Data

Authors (3):

Ivan Marin

Ankit Shukla

Sarang VK

View More author details

Table of Contents (11) Chapters

Big Data Analysis with Python

Preface

1. The Python Data Science Stack

2. Statistical Visualizations FREE CHAPTER

3. Working with Big Data Frameworks

4. Diving Deeper with Spark

5. Handling Missing Values and Correlation Analysis

6. Exploratory Data Analysis

7. Reproducibility in Big Data Analysis

8. Creating a Full Analysis Report

Appendix

Summary

In this chapter, we learned how to define a business problem from a data science perspective through a well-defined, structured approach. We started by understanding how to approach a business problem, how to gather the requirements from stakeholders and business experts, and how to define the business problem by developing an initial hypothesis.

Once the business problem was defined with data pipelines and workflows, we looked into understanding how to start the analysis on the gathered data in order to generate the KPIs and carry out descriptive analytics to identify the key trends and patterns in the historical data through various visualization techniques.

We also learned how a data science project life cycle is structured, from defining the business problem to various pre-processing techniques and model development. In the next chapter, we will be learning how to implement the concept of high reproducibility on a Jupyter notebook, and its importance in development.

The rest of the chapter is locked

Tech Concepts

Programming languages

Tech Tools

Unlimited access to the largest independent learning library in tech of over 8,000 expert-authored tech books and videos.

Innovative learning tools, including AI book assistants, code context explainers, and text-to-speech.

50+ new titles added per month and exclusive early access to books as they are being written.

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at $19.99/month. Cancel anytime

Authors (3)

Ivan Marin

Ivan Marin is a systems architect and data scientist working at Daitan Group, a Campinas-based software company. He designs big data systems for large volumes of data and implements machine learning pipelines end to end using Python and Spark. He is also an active organizer of data science, machine learning, and Python in So Paulo, and has given Python for data science courses at university level.

See other products by Ivan Marin

Ankit Shukla

Ankit Shukla is a data scientist working with World Wide Technology, a leading US-based technology solution provider, where he develops and deploys machine learning and artificial intelligence solutions to solve business problems and create actual dollar value for clients. He is also part of the company's R&D initiative, which is responsible for producing intellectual property, building capabilities in new areas, and publishing cutting-edge research in corporate white papers. Besides tinkering with AI/ML models, he likes to read and is a big-time foodie.

See other products by Ankit Shukla

Sarang VK

Sarang VK is a lead data scientist at StraitsBridge Advisors, where his responsibilities include requirement gathering, solutioning, development, and productization of scalable machine learning, artificial intelligence, and analytical solutions using open source technologies. Alongside this, he supports pre-sales and competency.

See other products by Sarang VK