You're reading from Data Engineering Best Practices Architect robust and cost-effective data solutions in the cloud era

Product type Paperback

Published in Oct 2024

Publisher Packt

ISBN-13 9781803244983

Length 550 pages

Edition 1st Edition

Languages

SQL

Tools

Google cloud SQL

Concepts

Data Engineering

Authors (2):

David Larochelle

Richard J. Schiller

View More author details

Table of Contents (21) Chapters

Preface

1. Chapter 1: Overview of the Business Problem Statement FREE CHAPTER

2. Chapter 2: A Data Engineer’s Journey – Background Challenges

3. Chapter 3: A Data Engineer’s Journey – IT’s Vision and Mission

4. Chapter 4: Architecture Principles

5. Chapter 5: Architecture Framework – Conceptual Architecture Best Practices

6. Chapter 6: Architecture Framework – Logical Architecture Best Practices

7. Chapter 7: Architecture Framework – Physical Architecture Best Practices

8. Chapter 8: Software Engineering Best Practice Considerations

9. Chapter 9: Key Considerations for Agile SDLC Best Practices

10. Chapter 10: Key Considerations for Quality Testing Best Practices

11. Chapter 11: Key Considerations for IT Operational Service Best Practices

12. Chapter 12: Key Considerations for Data Service Best Practices

13. Chapter 13: Key Considerations for Management Best Practices

14. Chapter 14: Key Considerations for Data Delivery Best Practices

15. Chapter 15: Other Considerations – Measures, Calculations, Restatements, and Data Science Best Practices

16. Chapter 16: Machine Learning Pipeline Best Practices and Processes

17. Chapter 17: Takeaway Summary – Putting It All Together

18. Chapter 18: Appendix and Use Cases

19. Index

Why subscribe?

20. Other Books You May Enjoy

Data annotation

So, you want to know how you, as a data engineer, can help a data scientist? You can begin by curating your data into indices that are semantically aligned with the object in your knowledge base. These, in turn, are correctly modeled for your business domains with the current known truths that are relevant for your enterprise. We bet you thought we would say, just build a vector store and make it available to your LLM, using the cloud provider of choice’s tool. However, you will fail to have that model propagate to production, since it would suffer from many of the failings of untuned GenAI models (such as hallucinations, quality errors, and the inappropriate exposure of training source text in model output). If you did as some of the cloud providers propose and built your RAG with an embedding, without semantic structure, you would be creating several iterations in your factory. Your vision should be to implement a knowledge-aware semantic form for embeddings...

The rest of the chapter is locked

You're reading from Data Engineering Best Practices Architect robust and cost-effective data solutions in the cloud era

Table of Contents (21) Chapters

Data annotation

Unlock this book and the full library FREE for 7 days

Authors (2)

Personalised recommendations for you