You're reading from Data Engineering with AWS - Second Edition

Product type Book

Published in Oct 2023

Publisher Packt

ISBN-13 9781804614426

Pages 636 pages

Edition 2nd Edition

Languages

Concepts

Data Engineering

Author (1):

Gareth Eagar

Table of Contents (24) Chapters

Preface

1. Section 1: AWS Data Engineering Concepts and Trends

2. An Introduction to Data Engineering

3. Data Management Architectures for Analytics

4. The AWS Data Engineer’s Toolkit

5. Data Governance, Security, and Cataloging

6. Section 2: Architecting and Implementing Data Engineering Pipelines and Transformations

7. Architecting Data Engineering Pipelines

8. Ingesting Batch and Streaming Data

9. Transforming Data to Optimize for Analytics

10. Identifying and Enabling Data Consumers

11. A Deeper Dive into Data Marts and Amazon Redshift

12. Orchestrating the Data Pipeline

13. Section 3: The Bigger Picture: Data Analytics, Data Visualization, and Machine Learning

14. Ad Hoc Queries with Amazon Athena

15. Visualizing Data with Amazon QuickSight

16. Enabling Artificial Intelligence and Machine Learning

17. Section 4: Modern Strategies: Open Table Formats, Data Mesh, DataOps, and Preparing for the Real World

18. Building Transactional Data Lakes

19. Implementing a Data Mesh Strategy

20. Building a Modern Data Platform on AWS

21. Wrapping Up the First Part of Your Learning Journey

22. Other Books You May Enjoy

23. Index

Examining examples of real-world data pipelines

The data pipeline examples that we have used in this book have been based on common types of transformations and pipelines, but they have been relatively simple examples. As you can imagine, in large organizations, the types of data pipelines that are built can be a lot more complex and may end up processing extremely large sets of data.

In this section, we will examine two examples of more complex data engineering pipelines from two very well-known organizations – Spotify and Netflix. Both of these companies have public blogs that cover software and data engineering, and the details provided about their pipelines in this section have been taken from the public information that’s been made available in a variety of blog posts and articles.

By learning more about these real-world big data pipelines, you can be better prepared for what to expect when you start working with very large datasets. Also, these examples...