You're reading from In-Memory Analytics with Apache Arrow Accelerate data analytics for efficient processing of flat and hierarchical data structures

Product type Paperback

Published in Sep 2024

Publisher Packt

ISBN-13 9781835461228

Length 406 pages

Edition 2nd Edition

Languages

Python

Tools

Apache arrow

Concepts

Data Engineering

Author (1):

Matthew Topol

View More author details

Table of Contents (18) Chapters

Preface

1. Part 1: Overview of What Arrow is, Its Capabilities, Benefits, and Goals

2. Chapter 1: Getting Started with Apache Arrow FREE CHAPTER

3. Chapter 2: Working with Key Arrow Specifications

4. Chapter 3: Format and Memory Handling

5. Part 2: Interoperability with Arrow: The Power of Open Standards

6. Chapter 4: Crossing the Language Barrier with the Arrow C Data API

7. Chapter 5: Acero: A Streaming Arrow Execution Engine

8. Chapter 6: Using the Arrow Datasets API

9. Chapter 7: Exploring Apache Arrow Flight RPC

10. Chapter 8: Understanding Arrow Database Connectivity (ADBC)

11. Chapter 9: Using Arrow with Machine Learning Workflows

12. Part 3: Real-World Examples, Use Cases, and Future Development

13. Chapter 10: Powered by Apache Arrow

14. Chapter 11: How to Leave Your Mark on Arrow

15. Chapter 12: Future Development and Plans

16. Index

Why subscribe?

17. Other Books You May Enjoy

Exploring Apache Arrow Flight RPC

Distributed systems have always interested me. A distributed system is like a really good puzzle; it’s immensely satisfying once you figure out how all the pieces fit together to achieve your goal. If you’re not familiar with the term, a distributed system is simply a situation where you have various components of a system spread across multiple machines on a network. The idea is to split up the work and coordinate efforts among the components to complete tasks more efficiently. A great example would be Apache Spark.

The goal of distributed systems is generally to provide a robust, scalable, and reliable conglomeration of components that efficiently perform operations by distributing work across a system. This often means large amounts of data flowing between various components so that the data can get processed, manipulated, or otherwise operated on. When it comes to Apache Arrow-formatted data, the Arrow project provides a remote...

The rest of the chapter is locked

You're reading from In-Memory Analytics with Apache Arrow Accelerate data analytics for efficient processing of flat and hierarchical data structures

Table of Contents (18) Chapters

Exploring Apache Arrow Flight RPC

Authors (1)

Personalised recommendations for you

You're reading from In-Memory Analytics with Apache Arrow Accelerate data analytics for efficient processing of flat and hierarchical data structures

Table of Contents (18) Chapters

Exploring Apache Arrow Flight RPC

Unlock this book and the full library FREE for 7 days

Authors (1)

Personalised recommendations for you