You're reading from Extending Power BI with Python and R Ingest, transform, enrich, and visualize data using the power of analytical languages

Product type Paperback

Published in Nov 2021

Publisher Packt

ISBN-13 9781801078207

Length 558 pages

Edition 1st Edition

Languages

Python

Tools

Power BI

Concepts

Business Intelligence

Author (1):

Luca Zavarella

View More author details

Table of Contents (22) Chapters

Preface

1. Section 1: Best Practices for Using R and Python in Power BI

2. Chapter 1: Where and How to Use R and Python Scripts in Power BI FREE CHAPTER

3. Chapter 2: Configuring R with Power BI

4. Chapter 3: Configuring Python with Power BI

5. Section 2: Data Ingestion and Transformation with R and Python in Power BI

6. Chapter 4: Importing Unhandled Data Objects

7. Chapter 5: Using Regular Expressions in Power BI

8. Chapter 6: Anonymizing and Pseudonymizing Your Data in Power BI

9. Chapter 7: Logging Data from Power BI to External Sources

10. Chapter 8: Loading Large Datasets beyond the Available RAM in Power BI

11. Section 3: Data Enrichment with R and Python in Power BI

12. Chapter 9: Calling External APIs to Enrich Your Data

13. Chapter 10: Calculating Columns Using Complex Algorithms

14. Chapter 11: Adding Statistics Insights: Associations

15. Chapter 12: Adding Statistics Insights: Outliers and Missing Values

16. Chapter 13: Using Machine Learning without Premium or Embedded Capacity

17. Section 3: Data Visualization with R in Power BI

18. Chapter 14: Exploratory Data Analysis

19. Chapter 15: Advanced Visualizations

20. Chapter 16: Interactive R Custom Visuals

21. Other Books You May Enjoy

Import large datasets with Python

In Chapter 3, Configuring Python with Power BI, we suggested that you install some of the most commonly used data management packages in your environment, including NumPy, pandas, and scikit-learn. The biggest limitation of these packages is that they cannot handle datasets larger than the RAM of the machine in which they are used, thus they are not able to scale to more than one machine. To comply with this limitation, distributed systems based on Spark, which has become a dominant tool in the big data analysis landscape, are often used. However, the move to these systems forces developers to have to rethink already-written code using an API called PySpark, born to use Spark objects with Python. This process is generally seen as causing delays in project delivery and causing frustration for developers, who master the libraries available for standard Python with much more confidence.

In response to the preceding issues, the community developed a...