Grouping and Statistics

data-science
Summarize categories and describe distributions.
  • Level: Beginner to intermediate
  • Estimated time: 35–50 minutes
  • You will learn: Summarize categories and describe distributions.
  • Practice in: Jupyter or Google Colab

Questions

  • What problem does Grouping and Statistics help us solve in a small Python program?
  • What should we predict before running the example?
  • What value, output, or error should we inspect after changing one line?

Objectives

  • Run a complete example for grouping and statistics in Colab.
  • Explain the example line by line using plain language.
  • Change one part of the code and predict the result before running it.
  • Recognize one common mistake and use the error message as evidence.

Hands-on episode: Grouping and Statistics

Groupby follows split-apply-combine: split rows by category, apply a summary such as mean or count, and combine the results into a new table.

We will learn this by running code, not by memorizing a definition first. Open the Colab notebook from the button above, find this section, and run each cell in order. Keep a small note beside the notebook with three columns: prediction, actual result, and what changed.

Example 1.1

Predict the mean score for each group before running.

import pandas as pd

data = pd.DataFrame({
"group": ["A", "A", "B"],
"score": [8, 10, 7],
})
print(data.groupby("group")["score"].agg(["count", "mean"]))

Run the cell once without editing it. If the result is different from your prediction, leave the prediction visible and write one sentence about the difference. That sentence is more useful than a perfect first guess.

Explain Example 1.1

  • Rows are split into group A and group B.
  • For each group, pandas counts scores and computes the mean.
  • The output combines one summary row per group.

Now explain the example out loud or in a Markdown cell. Use short sentences: “this line creates…”, “this name stores…”, “this output appears because…”. If you cannot explain a line yet, run only the lines above it and inspect the values that exist at that moment.

Challenge 1.1

NoteChallenge

Add another B score and see how both the count and mean change. The count is part of the evidence; a mean based on one row is less stable than a mean based on many rows.

Show a safe way to approach the challenge
  1. Copy Example 1.1 into a new Colab cell.
  2. Change exactly one value, name, condition, or line.
  3. Write the expected output before running the cell.
  4. Run the cell and compare the actual result with your prediction.
  5. If the result surprises you, undo the change and try a smaller one.

Suggested first move: Add another B score and see how both the count and mean change.

Debugging checkpoint 1.1

WarningDebugging checkpoint

Do not report an average without context. Check group sizes, spread, missing data, and whether the grouping variable actually answers your question. Averages can hide variation.

Do not debug by rewriting the whole example. Read the error type or surprising output, inspect the closest value with print(...) or type(...), then change one thing. This is the same routine you will use in larger projects.

Apply it

Create a DataFrame of scores by category. Compute count, mean, and median for each category, then write a two-sentence interpretation with one limitation.

Finish by adding a Markdown cell that answers: What did this example teach me that I can reuse in a project?

Key points

  • Learn the concept by running a complete, small example first.
  • Predict before execution so your thinking becomes visible.
  • Change one thing at a time so cause and effect stay clear.
  • Treat errors as clues about the exact line or value Python could not handle.

Why this matters

Summarize categories and describe distributions.

This lesson combines related subtopics that belong together in one learning conversation. You will still pause for a quiz after each section, but you do not need to jump between separate pages while building one clear explanation.

NoteGuiding questions

By the end of this lesson, you should be able to answer:

  • How do the sections in Grouping and Statistics fit together?
  • Which small example demonstrates each section?
  • Which debugging clue should I check first for each section?
NoteLearning objectives

You will practice how to:

  • explain the shared concept for this lesson;
  • use each section as one step in a larger workflow;
  • complete 2 short section quizzes before moving on;
  • connect examples, mistakes, and debugging routines.

Lesson map

  • 1. Grouping and Aggregation — Summarize data by categories with groupby.
  • 2. Descriptive Statistics — Summarize data with counts, averages, medians, spread, and distributions.

1. Grouping and Aggregation

Summarize data by categories with groupby.

TipAnalogy

groupby is sorting receipts into envelopes, totaling each envelope, then comparing totals.

What this means

Grouping splits data into categories, applies a calculation, and combines the results.

Example 1

Predict what will happen before you run the code.

import pandas as pd

df = pd.DataFrame({"group": ["a", "a", "b"], "score": [1, 3, 10]})
print(df.groupby("group")["score"].mean())

Step-by-step explanation

  1. import pandas as pd — pause here and say what this line reads, creates, changes, or displays.
  2. df = pd.DataFrame({"group": ["a", "a", "b"], "score": [1, 3, 10]}) — pause here and say what this line reads, creates, changes, or displays.
  3. print(df.groupby("group")["score"].mean()) — pause here and say what this line reads, creates, changes, or displays.

After running the example, compare the actual output with your prediction. If they differ, do not erase your prediction. The difference is the part that can teach you the most.

Challenge

NotePractice

Change one input value, predict the new output, run the code, and explain the difference in one sentence.

Show one possible solution path
  1. Copy Example 1 into Colab, Jupyter, or a .py file.
  2. Mark the line you plan to change.
  3. Write a one-sentence prediction.
  4. Run the changed code.
  5. If the result surprises you, restore the original and change a smaller part.

The goal is not to find the only correct answer. The goal is to create a small experiment where you can explain cause and effect.

Common mistakes

WarningCommon mistake

After groupby, check whether the result index and columns match what the next step expects.

When you get stuck, use this debugging routine:

  1. Read the last line of the error message or inspect the unexpected output.
  2. Find the smallest line of code that could be responsible.
  3. Print or inspect the value and type at that point.
  4. Change one thing.
  5. Run again and record what changed.

Check your understanding

This quiz checks the ideas in this section before you move on.

2. Descriptive Statistics

Summarize data with counts, averages, medians, spread, and distributions.

TipAnalogy

Descriptive statistics are a class photo summary: how many people, typical height, and variation, not every life story.

What this means

Descriptive statistics describe what is in the data without claiming causation.

Example 2

Predict what will happen before you run the code.

values = [2, 4, 4, 10]
mean = sum(values) / len(values)
print(mean)
print(max(values) - min(values))

Step-by-step explanation

  1. values = [2, 4, 4, 10] — pause here and say what this line reads, creates, changes, or displays.
  2. mean = sum(values) / len(values) — pause here and say what this line reads, creates, changes, or displays.
  3. print(mean) — pause here and say what this line reads, creates, changes, or displays.
  4. print(max(values) - min(values)) — pause here and say what this line reads, creates, changes, or displays.

After running the example, compare the actual output with your prediction. If they differ, do not erase your prediction. The difference is the part that can teach you the most.

Challenge

NotePractice

Change one input value, predict the new output, run the code, and explain the difference in one sentence.

Show one possible solution path
  1. Copy Example 1 into Colab, Jupyter, or a .py file.
  2. Mark the line you plan to change.
  3. Write a one-sentence prediction.
  4. Run the changed code.
  5. If the result surprises you, restore the original and change a smaller part.

The goal is not to find the only correct answer. The goal is to create a small experiment where you can explain cause and effect.

Common mistakes

WarningCommon mistake

The mean can be pulled by extreme values. Compare mean and median when data are skewed.

When you get stuck, use this debugging routine:

  1. Read the last line of the error message or inspect the unexpected output.
  2. Find the smallest line of code that could be responsible.
  3. Print or inspect the value and type at that point.
  4. Change one thing.
  5. Run again and record what changed.

Check your understanding

This quiz checks the ideas in this section before you move on.

Notebook and Colab practice

Open a blank notebook at https://colab.new. Use one section at a time: copy the Example 1, predict the result, run it, answer the section quiz, and then move to the next section. This is better than copying the entire page at once.

Instructor note

Teaching notes
  • Treat each section as a short teaching episode.
  • Pause for the section quiz before introducing the next section.
  • Ask learners to compare sections: what stayed the same, and what changed?
  • If time is short, teach the first two sections live and assign the rest as practice.

Key points

TipKey points
  • Grouping and Aggregation: Grouping splits data into categories, applies a calculation, and combines the results.
  • Descriptive Statistics: Descriptive statistics describe what is in the data without claiming causation.
  • Use the section quizzes as gates: review before moving on if a quiz feels uncertain.

References

  • pandas documentation: https://pandas.pydata.org/docs/
  • seaborn documentation: https://seaborn.pydata.org/
  • Matplotlib documentation: https://matplotlib.org/stable/
  • Python Tutorial: https://docs.python.org/3/tutorial/
  • Quarto OJS documentation: https://quarto.org/docs/interactive/ojs/
  • ipywidgets documentation: https://ipywidgets.readthedocs.io/en/stable/
Back to top