You're reading from Databricks Certified Associate Developer for Apache Spark Using Python The ultimate guide to getting certified in Apache Spark using practical examples with Python

Product type Paperback

Published in Jun 2024

Publisher Packt

ISBN-13 9781804619780

Length 274 pages

Edition 1st Edition

Languages

Python

Tools

Apache Spark

Concepts

Data Engineering

Author (1):

Saba Shah

View More author details

Table of Contents (18) Chapters

Preface

1. Part 1: Exam Overview

2. Chapter 1: Overview of the Certification Guide and Exam FREE CHAPTER

3. Part 2: Introducing Spark

4. Chapter 2: Understanding Apache Spark and Its Applications

5. Chapter 3: Spark Architecture and Transformations

6. Part 3: Spark Operations

7. Chapter 4: Spark DataFrames and their Operations

8. Chapter 5: Advanced Operations and Optimizations in Spark

9. Chapter 6: SQL Queries in Spark

10. Part 4: Spark Applications

11. Chapter 7: Structured Streaming in Spark

12. Chapter 8: Machine Learning with Spark ML

13. Part 5: Mock Papers

14. Chapter 9: Mock Test 1

15. Chapter 10: Mock Test 2

16. Index

Why subscribe?

17. Other Books You May Enjoy

What is Apache Spark?

Apache Spark is an open-source big data framework that is used for multiple big data applications. The strength of Spark lies in its superior parallel processing capabilities that makes it a leader in its domain.

According to its website (https://spark.apache.org/), “The most widely-used engine for scalable computing.”

The history of Apache Spark

Apache Spark started as a research project at the UC Berkeley AMPLab in 2009 and moved to an open source license in 2010. Later, in 2013, it came under the Apache Software Foundation (https://spark.apache.org/). It gained popularity after 2013, and today, it serves as a backbone for a large number of big data products across various Fortune 500 companies and has thousands of developers actively working on it.

Spark came into being because of limitations in the Hadoop MapReduce framework. MapReduce’s main premise was to read data from disk, distribute that data for parallel processing,...

The rest of the chapter is locked

You're reading from Databricks Certified Associate Developer for Apache Spark Using Python The ultimate guide to getting certified in Apache Spark using practical examples with Python

Table of Contents (18) Chapters

What is Apache Spark?

The history of Apache Spark

Authors (1)

Personalised recommendations for you

You're reading from Databricks Certified Associate Developer for Apache Spark Using Python The ultimate guide to getting certified in Apache Spark using practical examples with Python

Table of Contents (18) Chapters

What is Apache Spark?

The history of Apache Spark

Unlock this book and the full library FREE for 7 days

Authors (1)

Personalised recommendations for you