Batch Data Ingestion
In this chapter, we will look at the following key topics:
- Database migration using AWS DMS
- SaaS data ingestion using Amazon AppFlow
- Data ingestion using AWS Glue
- File and storage migration
So far, we have looked at creating scalable data lakes using Amazon S3 as the storage layer and AWS Glue Data Catalog as the metadata repository. We looked at how you can create layers of a data lake in S3 so that data can be systematically managed for specific personas in your organization. The very first layer we created in S3 was the raw layer, which is meant to store the source system data without any major changes. This also means that we need to first identify all the source systems that we need data from so that we can create a centralized data lake.
The mechanism by which we bring the data over into the raw layer of the data lake in S3 is also termed data ingestion. Data ingestion can either be in batches, where we bring the data over in...