You're reading from Data Engineering with Apache Spark, Delta Lake, and Lakehouse

Product type Book

Published in Oct 2021

Publisher Packt

ISBN-13 9781801077743

Pages 480 pages

Edition 1st Edition

Languages

Concepts

Data Processing

Author (1):

Manoj Kukreja

Table of Contents (17) Chapters

Preface

1. Section 1: Modern Data Engineering and Tools

2. Chapter 1: The Story of Data Engineering and Analytics

3. Chapter 2: Discovering Storage and Compute Data Lakes

4. Chapter 3: Data Engineering on Microsoft Azure

5. Section 2: Data Pipelines and Stages of Data Engineering

6. Chapter 4: Understanding Data Pipelines

7. Chapter 5: Data Collection Stage – The Bronze Layer

8. Chapter 6: Understanding Delta Lake

9. Chapter 7: Data Curation Stage – The Silver Layer

10. Chapter 8: Data Aggregation Stage – The Gold Layer

11. Section 3: Data Engineering Challenges and Effective Deployment Strategies

12. Chapter 9: Deploying and Monitoring Pipelines in Production

13. Chapter 10: Solving Data Engineering Challenges

14. Chapter 11: Infrastructure Provisioning

15. Chapter 12: Continuous Integration and Deployment (CI/CD) of Data Pipelines

16. Other Books You May Enjoy

Understanding Delta Lake

As mentioned before, modern data lakes lack some critical features such as ACID transactions, indexing, and versioning, which can negatively affect the reliability, quality, and performance of data.

Important Note

In data warehouse terms, ACID is the short form for atomicity, consistency, isolation, and durability of data. These properties are intended to make database transactions accurate, reliable, and permanent.

The following diagram depicts the properties:

Figure 6.2 – ACID properties in Delta Lake

Delta Lake functions as a layer on top of the distributed computing framework Apache Spark. By design, Apache Spark lacks the following key principles of transaction management:

It does not lock previous data during edit transactions, which means data may become unavailable during overwrites for a very brief period.
During data overwrites, there is a chance where the old data gets deleted yet the new data...

The rest of the chapter is locked

You're reading from Data Engineering with Apache Spark, Delta Lake, and Lakehouse

Table of Contents (17) Chapters

Understanding Delta Lake

Authors (1)

Personalised recommendations for you

You're reading from Data Engineering with Apache Spark, Delta Lake, and Lakehouse

Table of Contents (17) Chapters close

Understanding Delta Lake

Unlock this book and the full library FREE for 7 days

Authors (1)

Personalised recommendations for you

Table of Contents (17) Chapters