You're reading from Extending Power BI with Python and R Ingest, transform, enrich, and visualize data using the power of analytical languages

Product type Paperback

Published in Nov 2021

Publisher Packt

ISBN-13 9781801078207

Length 558 pages

Edition 1st Edition

Languages

Python

Tools

Power BI

Concepts

Business Intelligence

Author (1):

Luca Zavarella

View More author details

Table of Contents (22) Chapters

Preface

1. Section 1: Best Practices for Using R and Python in Power BI

2. Chapter 1: Where and How to Use R and Python Scripts in Power BI FREE CHAPTER

3. Chapter 2: Configuring R with Power BI

4. Chapter 3: Configuring Python with Power BI

5. Section 2: Data Ingestion and Transformation with R and Python in Power BI

6. Chapter 4: Importing Unhandled Data Objects

7. Chapter 5: Using Regular Expressions in Power BI

8. Chapter 6: Anonymizing and Pseudonymizing Your Data in Power BI

9. Chapter 7: Logging Data from Power BI to External Sources

10. Chapter 8: Loading Large Datasets beyond the Available RAM in Power BI

11. Section 3: Data Enrichment with R and Python in Power BI

12. Chapter 9: Calling External APIs to Enrich Your Data

13. Chapter 10: Calculating Columns Using Complex Algorithms

14. Chapter 11: Adding Statistics Insights: Associations

15. Chapter 12: Adding Statistics Insights: Outliers and Missing Values

16. Chapter 13: Using Machine Learning without Premium or Embedded Capacity

17. Section 3: Data Visualization with R in Power BI

18. Chapter 14: Exploratory Data Analysis

19. Chapter 15: Advanced Visualizations

20. Chapter 16: Interactive R Custom Visuals

21. Other Books You May Enjoy

Importing large datasets with R

The same scalability limitations illustrated for Python packages used to manipulate data also exist for R packages in the Tidyverse ecosystem. Even in R, it is not possible to use a dataset larger than the available RAM on the machine. The first solution that is adopted in these cases is also to switch to Spark-based distributed systems, which provide the SparkR language. It provides a distributed implementation of the DataFrame you are used to in R, supporting filtering, aggregation, and selection operations as you do with the dplyr package. For those of us who are fans of the Tidyverse world, RStudio actively develops the sparklyr package, which allows you to use all the functionality of dplyr, even for distributed DataFrames. However, adopting Spark-based systems to process CSVs that together take up little more than the RAM you have available on your machine may be overkill because of the overhead introduced by all the Java infrastructure needed...