Summary
If you are building up data pipelines and large systems, regardless of whether you are a data scientist or a software architect, you’re going to have to make a lot of decisions regarding which formats to use for various pieces of the system. You always want to choose the best format for the use case, and not just pick the latest trends and apply them everywhere. Many people hear about Arrow and either react by thinking that they need to use it everywhere for everything, or they wonder why we need yet another data format. The key takeaway I want you to understand is the differences in the problems that are trying to be solved.
If you need longer-term persistent storage either on disk or in the cloud, you typically want a storage format such as Parquet, ORC, or CSV, with the primary access cost being I/O time for these use cases, so you want to optimize to reduce that based on your access patterns. If you’re passing small messages around, such as metadata or...