Skip to content

Commit 1deb029

Browse files
Add materials for the Python statistics fundamentals tutorial
Python Statistics Fundamentals: How to Describe Your Data shipped without a companion folder. Its 2026 maintenance pass re-ran the tutorial on the current stack, so add the code here in runnable form. The tutorial teaches in the REPL. Each of its sections becomes a script that runs end to end, with the example data shared in datasets.py so a reader can swap in their own numbers in one place. Verified on the tutorial's pinned stack (Python 3.14.6, NumPy 2.5.3, SciPy 1.18.1, pandas 3.0.5, Matplotlib 3.11.2): all 13 scripts run, and the values match the tutorial, including the NumPy 2 scalar reprs and the linregress result. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 82a980d commit 1deb029

16 files changed

Lines changed: 854 additions & 0 deletions

‎python-statistics/README.md‎

Lines changed: 54 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,54 @@
1+
# Python Statistics Fundamentals: How to Describe Your Data
2+
3+
This folder contains supplementary code for the Real Python tutorial on [descriptive statistics in Python](https://realpython.com/python-statistics/). You can run the examples as they are, or continue your learning by experimenting with them on your own data.
4+
5+
The tutorial works through the examples in the REPL. Here, each section of the tutorial is collected into a script you can run end to end, so you can see every result at once and then change the numbers to see what happens.
6+
7+
## Setup
8+
9+
Create and activate a virtual environment, then install the requirements:
10+
11+
```bash
12+
$ python -m venv venv
13+
$ source venv/bin/activate
14+
(venv) $ python -m pip install -r requirements.txt -c constraints.txt
15+
```
16+
17+
The tutorial's examples were verified against the versions pinned in `constraints.txt`. Newer releases will usually work too, but NumPy 2 changed how scalars are displayed, so output from older versions won't match the tutorial exactly.
18+
19+
## Usage
20+
21+
After creating and activating your virtual environment, and installing the dependencies, you should be able to run each individual file normally:
22+
23+
```bash
24+
(venv) $ python central_tendency.py
25+
```
26+
27+
The statistics scripts print their results to the terminal. The plotting scripts open a Matplotlib window instead, one per figure — close each window to move on to the next.
28+
29+
| File | Tutorial section |
30+
| --- | --- |
31+
| `datasets.py` | The example data every other script imports |
32+
| `central_tendency.py` | Measures of Central Tendency |
33+
| `variability.py` | Measures of Variability |
34+
| `summary_statistics.py` | Summary of Descriptive Statistics |
35+
| `correlation.py` | Measures of Correlation Between Pairs of Data |
36+
| `two_dimensional_data.py` | Working With 2D Data: Axes |
37+
| `dataframes.py` | Working With 2D Data: DataFrames |
38+
| `plot_box.py` | Visualizing Data: Box Plots |
39+
| `plot_histogram.py` | Visualizing Data: Histograms |
40+
| `plot_pie.py` | Visualizing Data: Pie Charts |
41+
| `plot_bar.py` | Visualizing Data: Bar Charts |
42+
| `plot_xy.py` | Visualizing Data: X-Y Plots |
43+
| `plot_heatmap.py` | Visualizing Data: Heatmaps |
44+
45+
Each script is organized into functions named after the statistic they calculate, and running a file calls them in the order the tutorial introduces them. To work with just one measure, import it instead:
46+
47+
```pycon
48+
>>> from central_tendency import harmonic_mean
49+
>>> harmonic_mean()
50+
```
51+
52+
The example data lives in `datasets.py`, so you can swap in your own numbers in one place and re-run any script against them.
53+
54+
You can find more information and context on the code in [Python Statistics Fundamentals: How to Describe Your Data](https://realpython.com/python-statistics/).
Lines changed: 154 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,154 @@
1+
"""Measures of central tendency: mean, median, and mode.
2+
3+
Covers the "Measures of Central Tendency" section of
4+
https://realpython.com/python-statistics/
5+
"""
6+
7+
import math
8+
import statistics
9+
10+
import numpy as np
11+
import pandas as pd
12+
import scipy.stats
13+
14+
from datasets import one_dimensional
15+
16+
17+
def mean():
18+
"""Calculate the arithmetic mean with pure Python, NumPy, and pandas."""
19+
x, x_with_nan, y, y_with_nan, z, z_with_nan = one_dimensional()
20+
21+
print("Pure Python:")
22+
print(f" sum(x) / len(x) = {sum(x) / len(x)}")
23+
print(f" statistics.mean(x) = {statistics.mean(x)}")
24+
print(f" statistics.fmean(x) = {statistics.fmean(x)}")
25+
26+
# A single nan in the data makes the result nan.
27+
print(f" statistics.mean(x_with_nan) = {statistics.mean(x_with_nan)}")
28+
29+
print("NumPy:")
30+
print(f" np.mean(y) = {np.mean(y)}")
31+
print(f" y.mean() = {y.mean()}")
32+
print(f" np.mean(y_with_nan) = {np.mean(y_with_nan)}")
33+
# Use the nan-safe variant to ignore missing values instead.
34+
print(f" np.nanmean(y_with_nan) = {np.nanmean(y_with_nan)}")
35+
36+
print("pandas:")
37+
# pandas skips nan values by default.
38+
print(f" z.mean() = {z.mean()}")
39+
print(f" z_with_nan.mean() = {z_with_nan.mean()}")
40+
41+
42+
def weighted_mean():
43+
"""Calculate the weighted mean, where each item has its own weight."""
44+
x = [8.0, 1, 2.5, 4, 28.0]
45+
w = [0.1, 0.2, 0.3, 0.25, 0.15]
46+
47+
wmean = sum(w_ * x_ for (x_, w_) in zip(x, w)) / sum(w)
48+
print(f"Pure Python: {wmean}")
49+
50+
y, z, w_array = np.array(x), pd.Series(x), np.array(w)
51+
print(f"np.average(y, weights=w) = {np.average(y, weights=w_array)}")
52+
print(f"np.average(z, weights=w) = {np.average(z, weights=w_array)}")
53+
print(f"(w * y).sum() / w.sum() = {(w_array * y).sum() / w_array.sum()}")
54+
55+
56+
def harmonic_mean():
57+
"""Calculate the harmonic mean, the reciprocal of the mean reciprocal."""
58+
x, _, y, _, z, _ = one_dimensional()
59+
60+
hmean = len(x) / sum(1 / item for item in x)
61+
print(f"Pure Python: {hmean}")
62+
print(f"statistics.harmonic_mean(x): {statistics.harmonic_mean(x)}")
63+
print(f"scipy.stats.hmean(y): {scipy.stats.hmean(y)}")
64+
print(f"scipy.stats.hmean(z): {scipy.stats.hmean(z)}")
65+
66+
# A negative value has no harmonic mean.
67+
try:
68+
statistics.harmonic_mean([1, 2, -2])
69+
except statistics.StatisticsError as error:
70+
print(f"harmonic_mean([1, 2, -2]) raises StatisticsError: {error}")
71+
72+
73+
def geometric_mean():
74+
"""Calculate the geometric mean, the nth root of the product."""
75+
x, _, y, _, z, _ = one_dimensional()
76+
77+
gmean = 1
78+
for item in x:
79+
gmean *= item
80+
gmean **= 1 / len(x)
81+
82+
print(f"Pure Python: {gmean}")
83+
print(f"statistics.geometric_mean(x): {statistics.geometric_mean(x)}")
84+
print(f"scipy.stats.gmean(y): {scipy.stats.gmean(y)}")
85+
print(f"scipy.stats.gmean(z): {scipy.stats.gmean(z)}")
86+
87+
88+
def median():
89+
"""Find the median, the middle value of the sorted data."""
90+
x, x_with_nan, y, y_with_nan, z, z_with_nan = one_dimensional()
91+
92+
n = len(x)
93+
if n % 2:
94+
median_ = sorted(x)[round(0.5 * (n - 1))]
95+
else:
96+
x_ord, index = sorted(x), round(0.5 * n)
97+
median_ = 0.5 * (x_ord[index - 1] + x_ord[index])
98+
print(f"Pure Python: {median_}")
99+
100+
print(f"statistics.median(x): {statistics.median(x)}")
101+
# With an even number of items, the median averages the two middle values.
102+
print(f"median(x[:-1]): {statistics.median(x[:-1])}")
103+
print(f"median_low(x[:-1]): {statistics.median_low(x[:-1])}")
104+
print(f"median_high(x[:-1]): {statistics.median_high(x[:-1])}")
105+
106+
print(f"np.median(y): {np.median(y)}")
107+
print(f"np.nanmedian(y_with_nan): {np.nanmedian(y_with_nan)}")
108+
print(f"z.median(): {z.median()}")
109+
print(f"z_with_nan.median(): {z_with_nan.median()}")
110+
111+
112+
def mode():
113+
"""Find the mode, the value that appears most often."""
114+
u = [2, 3, 2, 8, 12]
115+
v = [12, 15, 12, 15, 21, 15, 12]
116+
117+
mode_ = max((u.count(item), item) for item in set(u))[1]
118+
print(f"Pure Python: {mode_}")
119+
print(f"statistics.mode(u): {statistics.mode(u)}")
120+
print(f"statistics.multimode(u): {statistics.multimode(u)}")
121+
122+
# multimode() returns every modal value, mode() only the first.
123+
print(f"statistics.mode(v): {statistics.mode(v)}")
124+
print(f"statistics.multimode(v): {statistics.multimode(v)}")
125+
126+
u_array, v_array = np.array(u), np.array(v)
127+
print(f"scipy.stats.mode(u): {scipy.stats.mode(u_array)}")
128+
129+
result = scipy.stats.mode(v_array)
130+
print(f"scipy.stats.mode(v): {result}")
131+
print(f" .mode = {result.mode}")
132+
print(f" .count = {result.count}")
133+
134+
# pandas returns a Series, so it can report several modal values at once.
135+
u_series = pd.Series(u)
136+
v_series = pd.Series(v)
137+
w_series = pd.Series([2, 2, math.nan])
138+
print(f"u.mode():\n{u_series.mode()}")
139+
print(f"v.mode():\n{v_series.mode()}")
140+
print(f"w.mode():\n{w_series.mode()}")
141+
142+
143+
if __name__ == "__main__":
144+
for section in (
145+
mean,
146+
weighted_mean,
147+
harmonic_mean,
148+
geometric_mean,
149+
median,
150+
mode,
151+
):
152+
title = section.__name__.replace("_", " ").title()
153+
print(f"\n{title}\n{'-' * len(title)}")
154+
section()

‎python-statistics/constraints.txt‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,4 @@
1+
numpy==2.5.3
2+
scipy==1.18.1
3+
pandas==3.0.5
4+
matplotlib==3.11.2

‎python-statistics/correlation.py‎

Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,71 @@
1+
"""Measures of correlation between pairs of data.
2+
3+
Covers the "Measures of Correlation Between Pairs of Data" section of
4+
https://realpython.com/python-statistics/
5+
"""
6+
7+
import numpy as np
8+
import scipy.stats
9+
10+
from datasets import correlation_data
11+
12+
13+
def covariance():
14+
"""Calculate the covariance, which measures joint variability."""
15+
x, y, x_, y_, x__, y__ = correlation_data()
16+
17+
n = len(x)
18+
mean_x, mean_y = sum(x) / n, sum(y) / n
19+
cov_xy = sum((x[k] - mean_x) * (y[k] - mean_y) for k in range(n)) / (n - 1)
20+
print(f"Pure Python: {cov_xy}")
21+
22+
# np.cov() returns the covariance matrix, not a single number. Its
23+
# diagonal holds the variances and its off-diagonal the covariance.
24+
cov_matrix = np.cov(x_, y_)
25+
print(f"np.cov(x_, y_) =\n{cov_matrix}")
26+
print(f" variance of x: {x_.var(ddof=1)}")
27+
print(f" variance of y: {y_.var(ddof=1)}")
28+
print(f" covariance: {cov_matrix[0, 1]}")
29+
30+
print(f"x__.cov(y__): {x__.cov(y__)}")
31+
32+
33+
def correlation_coefficient():
34+
"""Calculate Pearson's r, the normalized measure of correlation."""
35+
x, y, x_, y_, x__, y__ = correlation_data()
36+
37+
n = len(x)
38+
mean_x, mean_y = sum(x) / n, sum(y) / n
39+
var_x = sum((item - mean_x) ** 2 for item in x) / (n - 1)
40+
var_y = sum((item - mean_y) ** 2 for item in y) / (n - 1)
41+
std_x, std_y = var_x**0.5, var_y**0.5
42+
cov_xy = sum((x[k] - mean_x) * (y[k] - mean_y) for k in range(n)) / (n - 1)
43+
print(f"Pure Python: {cov_xy / (std_x * std_y)}")
44+
45+
# pearsonr() returns the coefficient and the p-value together.
46+
r, p = scipy.stats.pearsonr(x_, y_)
47+
print(f"scipy.stats.pearsonr(): r={r}, p={p}")
48+
49+
corr_matrix = np.corrcoef(x_, y_)
50+
print(f"np.corrcoef(x_, y_) =\n{corr_matrix}")
51+
print(f" r = {corr_matrix[0, 1]}")
52+
53+
print(f"x__.corr(y__): {x__.corr(y__)}")
54+
55+
56+
def linear_regression():
57+
"""Fit a regression line, which also reports the correlation."""
58+
*_, x_, y_, __, ___ = correlation_data()
59+
60+
result = scipy.stats.linregress(x_, y_)
61+
print(result)
62+
print(f" slope = {result.slope}")
63+
print(f" intercept = {result.intercept}")
64+
print(f" rvalue = {result.rvalue}")
65+
66+
67+
if __name__ == "__main__":
68+
for section in (covariance, correlation_coefficient, linear_regression):
69+
title = section.__name__.replace("_", " ").title()
70+
print(f"\n{title}\n{'-' * len(title)}")
71+
section()

‎python-statistics/dataframes.py‎

Lines changed: 64 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,64 @@
1+
"""Statistics with pandas DataFrame objects.
2+
3+
Covers the "DataFrames" section of https://realpython.com/python-statistics/
4+
"""
5+
6+
import pandas as pd
7+
8+
from datasets import two_dimensional
9+
10+
11+
def build_dataframe():
12+
"""Return the labeled DataFrame used throughout this script."""
13+
row_names = ["first", "second", "third", "fourth", "fifth"]
14+
col_names = ["A", "B", "C"]
15+
return pd.DataFrame(two_dimensional(), index=row_names, columns=col_names)
16+
17+
18+
def by_column_and_row():
19+
"""Calculate statistics across columns and across rows."""
20+
df = build_dataframe()
21+
print(f"df =\n{df}\n")
22+
23+
# DataFrame methods default to working column by column.
24+
print(f"df.mean():\n{df.mean()}\n")
25+
print(f"df.var():\n{df.var()}\n")
26+
27+
# Pass axis=1 to work row by row instead.
28+
print(f"df.mean(axis=1):\n{df.mean(axis=1)}\n")
29+
print(f"df.var(axis=1):\n{df.var(axis=1)}")
30+
31+
32+
def single_column():
33+
"""Pull one column out as a Series and work with it directly."""
34+
df = build_dataframe()
35+
36+
print(f"df['A']:\n{df['A']}\n")
37+
print(f"df['A'].mean(): {df['A'].mean()}")
38+
print(f"df['A'].var(): {df['A'].var()}")
39+
40+
41+
def to_numpy():
42+
"""Drop the labels and get back a plain NumPy array."""
43+
df = build_dataframe()
44+
45+
print(f"df.values:\n{df.values}\n")
46+
print(f"df.to_numpy():\n{df.to_numpy()}")
47+
48+
49+
def describe():
50+
"""Summarize every column at once."""
51+
df = build_dataframe()
52+
53+
print(f"df.describe():\n{df.describe()}\n")
54+
55+
# The summary is a DataFrame too, so .at[] reaches a single value.
56+
print(f"df.describe().at['mean', 'A']: {df.describe().at['mean', 'A']}")
57+
print(f"df.describe().at['50%', 'B']: {df.describe().at['50%', 'B']}")
58+
59+
60+
if __name__ == "__main__":
61+
for section in (by_column_and_row, single_column, to_numpy, describe):
62+
title = section.__name__.replace("_", " ").title()
63+
print(f"\n{title}\n{'-' * len(title)}")
64+
section()

0 commit comments

Comments
 (0)