You're reading from Data Engineering Best Practices Architect robust and cost-effective data solutions in the cloud era

Product type Paperback

Published in Oct 2024

Publisher Packt

ISBN-13 9781803244983

Length 550 pages

Edition 1st Edition

Languages

SQL

Tools

Google cloud SQL

Concepts

Data Engineering

Authors (2):

David Larochelle

Richard J. Schiller

View More author details

Table of Contents (21) Chapters

Preface

1. Chapter 1: Overview of the Business Problem Statement FREE CHAPTER

2. Chapter 2: A Data Engineer’s Journey – Background Challenges

3. Chapter 3: A Data Engineer’s Journey – IT’s Vision and Mission

4. Chapter 4: Architecture Principles

5. Chapter 5: Architecture Framework – Conceptual Architecture Best Practices

6. Chapter 6: Architecture Framework – Logical Architecture Best Practices

7. Chapter 7: Architecture Framework – Physical Architecture Best Practices

8. Chapter 8: Software Engineering Best Practice Considerations

9. Chapter 9: Key Considerations for Agile SDLC Best Practices

10. Chapter 10: Key Considerations for Quality Testing Best Practices

11. Chapter 11: Key Considerations for IT Operational Service Best Practices

12. Chapter 12: Key Considerations for Data Service Best Practices

13. Chapter 13: Key Considerations for Management Best Practices

14. Chapter 14: Key Considerations for Data Delivery Best Practices

15. Chapter 15: Other Considerations – Measures, Calculations, Restatements, and Data Science Best Practices

16. Chapter 16: Machine Learning Pipeline Best Practices and Processes

17. Chapter 17: Takeaway Summary – Putting It All Together

18. Chapter 18: Appendix and Use Cases

19. Index

Why subscribe?

20. Other Books You May Enjoy

Summary

This chapter addressed important data delivery practices to be considered as part of your data-engineered solution. When starting out, we stated that the data consumers’ data processing use cases always require data to be consistent, complete, and semantically correct. This chapter elaborated on that goal with best practices and examples of choices that you have to make as you implement your data solutions.

You were exposed to data streaming considerations and how bulk operations can be treated as streaming operations of smaller bulk (or micro-batch) sizes. This enables you to tune the entire system for the best performance given the technologies being used. Best practices for publishing and subscribing to data were outlined and then elaborated upon. Data flow was discussed with a study of Google’s implementation in detail with Apache Beam.

We also explored how best to organize huge volumes of data while making the output of the data factory available for...