You're reading from Databricks Certified Associate Developer for Apache Spark Using Python The ultimate guide to getting certified in Apache Spark using practical examples with Python

Product type Paperback

Published in Jun 2024

Publisher Packt

ISBN-13 9781804619780

Length 274 pages

Edition 1st Edition

Languages

Python

Tools

Apache Spark

Concepts

Data Engineering

Author (1):

Saba Shah

View More author details

Table of Contents (18) Chapters

Preface

1. Part 1: Exam Overview

2. Chapter 1: Overview of the Certification Guide and Exam FREE CHAPTER

3. Part 2: Introducing Spark

4. Chapter 2: Understanding Apache Spark and Its Applications

5. Chapter 3: Spark Architecture and Transformations

6. Part 3: Spark Operations

7. Chapter 4: Spark DataFrames and their Operations

8. Chapter 5: Advanced Operations and Optimizations in Spark

9. Chapter 6: SQL Queries in Spark

10. Part 4: Spark Applications

11. Chapter 7: Structured Streaming in Spark

12. Chapter 8: Machine Learning with Spark ML

13. Part 5: Mock Papers

14. Chapter 9: Mock Test 1

15. Chapter 10: Mock Test 2

16. Index

Why subscribe?

17. Other Books You May Enjoy

Creating DataFrame operations

As we have already discussed, DataFrames are the main building blocks of Spark data. They consist of rows and column data structures.

DataFrames in PySpark are created using the pyspark.sql.SparkSession.createDataFrame function. You can use lists, lists of lists, tuples, dictionaries, Pandas DataFrames, RDDs, and pyspark.sql.Rows to create DataFrames.

Spark DataFrames also has an argument named schema that specifies the schema of the DataFrame. You can either choose to specify the schema explicitly or let Spark infer the schema from the DataFrame itself. If you don’t specify this argument in the code, Spark will infer the schema on its own.

There are different ways to create DataFrames in Spark. Some of them are explained in the following sections.

Using a list of rows

The first way to create DataFrames we see is by using rows of data. You can think of rows of data as lists. They would share common header values for each of the values...

The rest of the chapter is locked

You're reading from Databricks Certified Associate Developer for Apache Spark Using Python The ultimate guide to getting certified in Apache Spark using practical examples with Python

Table of Contents (18) Chapters

Creating DataFrame operations

Using a list of rows

Authors (1)

Personalised recommendations for you

You're reading from Databricks Certified Associate Developer for Apache Spark Using Python The ultimate guide to getting certified in Apache Spark using practical examples with Python

Table of Contents (18) Chapters

Creating DataFrame operations

Using a list of rows

Unlock this book and the full library FREE for 7 days

Authors (1)

Personalised recommendations for you