You're reading from Data Engineering with Python Work with massive datasets to design data models and automate data pipelines using Python

Product type Paperback

Published in Oct 2020

Publisher Packt

ISBN-13 9781839214189

Length 356 pages

Edition 1st Edition

Languages

Python

Concepts

Data Analysis

Author (1):

Paul Crickard

View More author details

Table of Contents (21) Chapters

Preface

1. Section 1: Building Data Pipelines – Extract Transform, and Load

2. Chapter 1: What is Data Engineering? FREE CHAPTER

3. Chapter 2: Building Our Data Engineering Infrastructure

4. Chapter 3: Reading and Writing Files

5. Chapter 4: Working with Databases

6. Chapter 5: Cleaning, Transforming, and Enriching Data

7. Chapter 6: Building a 311 Data Pipeline

8. Section 2:Deploying Data Pipelines in Production

9. Chapter 7: Features of a Production Pipeline

10. Chapter 8: Version Control with the NiFi Registry

11. Chapter 9: Monitoring Data Pipelines

12. Chapter 10: Deploying Data Pipelines

13. Chapter 11: Building a Production Data Pipeline

14. Section 3:Beyond Batch – Building Real-Time Data Pipelines

15. Chapter 12: Building a Kafka Cluster

16. Chapter 13: Streaming Data with Apache Kafka

17. Chapter 14: Data Processing with Apache Spark

18. Chapter 15: Real-Time Edge Data with MiNiFi, Kafka, and Spark

19. Other Books You May Enjoy

Leave a review - let other readers know what you think

Appendix

Summary

In this chapter, you learned how MiNiFi provides a means by which you can stream data to a NiFi instance. With MiNiFi, you can capture data from sensors, smaller devices such as a Raspberry Pi, or on regular servers where the data lives, without needing a full NiFi install. You learned how to set up and configure a remote processor group that allows you to talk to a remote NiFi instance.

In the Appendix, you will learn how you can cluster NiFi to run your data pipelines on different machines so that you can further distribute the load. This will allow you to reserve servers for specific tasks, or to spread large amounts of data horizontally across the cluster. By combining NiFi, Kafka, and Spark into clusters, you will be able to process more data than any single machine.

The rest of the chapter is locked

You're reading from Data Engineering with Python Work with massive datasets to design data models and automate data pipelines using Python

Table of Contents (21) Chapters

Summary

Authors (1)

Other recommended products

Personalised recommendations for you

You're reading from Data Engineering with Python Work with massive datasets to design data models and automate data pipelines using Python

Table of Contents (21) Chapters

Summary

Unlock this book and the full library FREE for 7 days

Authors (1)

Other recommended products

Personalised recommendations for you