You're reading from Data Engineering with Apache Spark, Delta Lake, and Lakehouse Create scalable pipelines that ingest, curate, and aggregate complex data in a timely and secure way

Product type Paperback

Published in Oct 2021

Publisher Packt

ISBN-13 9781801077743

Length 480 pages

Edition 1st Edition

Languages

Python

Tools

Apache Spark

Concepts

Data Engineering

Author (1):

Manoj Kukreja

View More author details

Table of Contents (17) Chapters

Preface

1. Section 1: Modern Data Engineering and Tools

2. Chapter 1: The Story of Data Engineering and Analytics FREE CHAPTER

3. Chapter 2: Discovering Storage and Compute Data Lakes

4. Chapter 3: Data Engineering on Microsoft Azure

5. Section 2: Data Pipelines and Stages of Data Engineering

6. Chapter 4: Understanding Data Pipelines

7. Chapter 5: Data Collection Stage – The Bronze Layer

8. Chapter 6: Understanding Delta Lake

9. Chapter 7: Data Curation Stage – The Silver Layer

10. Chapter 8: Data Aggregation Stage – The Gold Layer

11. Section 3: Data Engineering Challenges and Effective Deployment Strategies

12. Chapter 9: Deploying and Monitoring Pipelines in Production

13. Chapter 10: Solving Data Engineering Challenges

14. Chapter 11: Infrastructure Provisioning

15. Chapter 12: Continuous Integration and Deployment (CI/CD) of Data Pipelines

16. Other Books You May Enjoy

Chapter 6: Understanding Delta Lake

In the previous chapter, we created the bronze layer of the lakehouse. The bronze layer stores raw data in the native form as collected from the data sources. The problem is that raw data is not in a shape that can be readily consumed for analytical operations.

As a data engineer, it is your responsibility to convert raw data into a shape and form that becomes ready for use analytical workloads. In this chapter, we will further advance our learning to cleanse raw data. The process of cleansing data involves applying the logic that cleans and standardizes data followed by writing it to the silver layer of the lakehouse.

But that is not all – the silver layer should store data in an open format that supports ACID (atomicity, consistency, isolation, and durability) transactions. This is done by using the Delta Lake engine. Before we start building the silver layer, we need to completely understand some critical features of Delta Lake and...

The rest of the chapter is locked

You're reading from Data Engineering with Apache Spark, Delta Lake, and Lakehouse Create scalable pipelines that ingest, curate, and aggregate complex data in a timely and secure way

Table of Contents (17) Chapters

Chapter 6: Understanding Delta Lake

Authors (1)

Personalised recommendations for you

You're reading from Data Engineering with Apache Spark, Delta Lake, and Lakehouse Create scalable pipelines that ingest, curate, and aggregate complex data in a timely and secure way

Table of Contents (17) Chapters

Chapter 6: Understanding Delta Lake

Unlock this book and the full library FREE for 7 days

Authors (1)

Personalised recommendations for you