You're reading from Python Data Cleaning Cookbook Modern techniques and Python tools to detect and remove dirty data and extract key insights

Product type Paperback

Published in Dec 2020

Publisher Packt

ISBN-13 9781800565661

Length 436 pages

Edition 1st Edition

Languages

Python

Tools

Pandas

Concepts

Data Analysis

Authors (2):

Michael Walker

Michael B Walker

View More author details

Table of Contents (12) Chapters

Preface

1. Chapter 1: Anticipating Data Cleaning Issues when Importing Tabular Data into pandas

2. Chapter 2: Anticipating Data Cleaning Issues when Importing HTML and JSON into pandas FREE CHAPTER

3. Chapter 3: Taking the Measure of Your Data

4. Chapter 4: Identifying Missing Values and Outliers in Subsets of Data

5. Chapter 5: Using Visualizations for the Identification of Unexpected Values

6. Chapter 6: Cleaning and Exploring Data with Series Operations

7. Chapter 7: Fixing Messy Data when Aggregating

8. Chapter 8: Addressing Data Issues When Combining DataFrames

9. Chapter 9: Tidying and Reshaping Data

10. Chapter 10: User-Defined Functions and Classes to Automate Data Cleaning

11. Other Books You May Enjoy

Leave a review - let other readers know what you think

Using more complicated aggregation functions with groupby

In the previous recipe, we created a groupby DataFrame object and used it to run summary statistics by groups. We use chaining in this recipe to create the groups, choose the aggregation variable(s), and select the aggregation function(s), all in one line. We also take advantage of the flexibility of the groupby object, which allows us to choose the aggregation columns and functions in a variety of ways.

Getting ready

We will work with the National Longitudinal Survey of Youth (NLS) data in this recipe.

Data note

The NLS, administered by the United States Bureau of Labor Statistics, are longitudinal surveys of individuals who were in high school in 1997 when the surveys started. Participants were surveyed each year through 2018. The surveys are available for public use at nlsinfo.org.

How to do it…

We do more complicated aggregations with groupby than we did in the previous recipe, taking advantage...

The rest of the chapter is locked

You're reading from Python Data Cleaning Cookbook Modern techniques and Python tools to detect and remove dirty data and extract key insights

Table of Contents (12) Chapters

Using more complicated aggregation functions with groupby

Getting ready

How to do it…

Authors (2)

Other recommended products

Personalised recommendations for you

You're reading from Python Data Cleaning Cookbook Modern techniques and Python tools to detect and remove dirty data and extract key insights

Table of Contents (12) Chapters

Using more complicated aggregation functions with groupby

Getting ready

How to do it…

Unlock this book and the full library FREE for 7 days

Authors (2)

Other recommended products

Personalised recommendations for you