Skip to content
Léonel Vodounou
Back to the blog

Foundations

Descriptive Statistics

Before diving into machine learning algorithms, we need a solid grounding in descriptive statistics, the tools that let us summarize and understand data at a glance.

Léonel VODOUNOU

March 26, 2026 · 14 min read

Introduction

To build models that perform well, you need an intimate understanding of the nature of your data. Descriptive statistics give us exactly the mathematical concepts we need to analyze and summarize any dataset.


What Are Descriptive Statistics?

Descriptive statistics are the tools we use to summarize and describe data. Instead of looking at thousands of numbers, we extract a few key indicators that tell the essential story.

For example: “This app has a 4.2-star rating based on 10,000 reviews.” That single number condenses thousands of individual opinions into one useful piece of information.

In this guide, we will cover the fundamentals: the mean, the median and the mode (where is the “center” of your data?), as well as the variance and the standard deviation (how spread out is your data?).

These concepts show up everywhere: from making sense of survey results to building machine learning models, by way of analyzing probability distributions.


Measures of Central Tendency

The Mean

The mean is the most widely used measure of central tendency. It represents the balance point of a dataset.

Definition

The arithmetic mean is the sum of all values divided by the number of values.

xˉ=1n∑i=1nxi=x1+x2+⋯+xnn\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i = \frac{x_1 + x_2 + \dots + x_n}{n}

Example: Stable temperatures

You measure the temperature of a server room over 5 days: 20°C, 21°C, 20°C, 22°C, 22°C.

  • Mean: (20+21+20+22+22)/5=21(20 + 21 + 20 + 22 + 22) / 5 = 21 °C.

Here, the mean describes the situation perfectly.


The Median (Middle Value)

The median is the value that splits the dataset into two equal groups: 50% of the values lie below it and 50% lie above it. It is favored for its robustness.

Definition

Once the data is sorted in ascending order:

Median={xn+12if n is oddxn2+xn2+12if n is even\text{Median} = \begin{cases} x_{\frac{n+1}{2}} & \text{if } n \text{ is odd} \\ \frac{x_{\frac{n}{2}} + x_{\frac{n}{2} + 1}}{2} & \text{if } n \text{ is even} \end{cases}

Example: The impact of an outlier

Consider the monthly incomes of 5 people.

  • Case A (Homogeneous distribution): €1,500, €1,600, €1,700, €1,800, €1,900

    • Mean: €1,700
    • Median: €1,700 (The two measures agree).
  • Case B (With an extreme outlier): €1,500, €1,600, €1,700, €1,800, €100,000

    • Median: €1,700 (It stays stable and still represents the group).
    • Mean: €21,320 (It explodes and no longer represents anyone in the group).

The Mode (Most Frequent)

The mode is the most frequent value. It is the only measure of central tendency that works for qualitative (categorical) data.

Definition

The mode is the value with the highest count. A distribution can be:

  • Unimodal (a single mode).
  • Bimodal (two modes, often a sign of two distinct subpopulations).
  • Multimodal (more than two modes).

Example: Industrial maintenance

We record the cause of each failure on a production line: Electrical, Mechanical, Electrical, Electrical, Software, Mechanical, Electrical.

  • Mode: Electrical (4 occurrences).

That is the key piece of information for prioritizing repairs!


Empirical Quantiles

Empirical quantiles are values that divide the ordered sample into a given number of parts of equal size. They generalize the median.

  • With 2 parts, we recover the empirical median: x~n\tilde{x}_n.
  • With 4 parts, we speak of quartiles, written q~n,1/4\tilde{q}_{n, 1/4}, q~n,1/2\tilde{q}_{n, 1/2} and q~n,3/4\tilde{q}_{n, 3/4}. Naturally, q~n,1/2=x~n\tilde{q}_{n, 1/2} = \tilde{x}_n.
  • With 10 parts, we speak of deciles, written q~n,1/10,…,q~n,9/10\tilde{q}_{n, 1/10}, \dots, \tilde{q}_{n, 9/10}.
  • With 100 parts, we speak of percentiles, written q~n,1/100,…,q~n,99/100\tilde{q}_{n, 1/100}, \dots, \tilde{q}_{n, 99/100}.

Mathematical Definition

For a sample of size nn, the empirical quantile of order pp (where p∈]0,1[p \in ]0, 1[) is computed from the data sorted in ascending order, written (x1∗,x2∗,…,xn∗)(x^*_1, x^*_2, \dots, x^*_n).

It is defined as follows:

q~n,p={12(xnp∗+xnp+1∗)if np is an integerx⌊np⌋+1∗otherwise\tilde{q}_{n,p} = \begin{cases} \frac{1}{2}(x^*_{np} + x^*_{np+1}) & \text{if } np \text{ is an integer} \\ x^*_{\lfloor np \rfloor + 1} & \text{otherwise} \end{cases}

Note: ⌊np⌋\lfloor np \rfloor is the integer part of n×pn \times p. For example, if n=10n=10 and we are looking for the first quartile (p=0.25p=0.25), then np=2.5np = 2.5. Since 2.52.5 is not an integer, the quantile is the value of rank ⌊2.5⌋+1=3\lfloor 2.5 \rfloor + 1 = 3, that is, the third value of the sorted sample.


Why Is It Useful?

Quantile analysis is fundamental to understanding the structure of data. For example, comparing the 9th decile (D9D_9) with the 1st decile (D1D_1) is a standard way to measure income inequality in a population, where a simple mean would hide the real gaps.


Measures of Shape

Skewness

So far, we have studied three ways of finding the center of your data: the mean, the median and the mode. But here is a question: what if the mean and the median give different results? This happens when your data is not symmetric.

Skewness describes the shape of your data’s distribution.

Refresher: what is a distribution?

Imagine you work in Human Resources and are in charge of a survey on how pay is structured across your company.

To represent this data, you use a horizontal axis divided into salary bands. For each employee, you stack one unit (an elementary brick) in the matching column.

Together, these columns form what is called a histogram. The curve that follows the top of the columns defines the distribution of your data. In statistics, this distribution is the complete picture of how the observations are spread across the variable being studied. It lets you see at once where observations are dense and which values are most frequent.

1 brick = 1 employee Distribution 20 30 40 50 60 70 80 90 100 110 120 130 0510 Annual gross salary (k€) Number of employees
Figure 1: Annual salaries of 60 employees at a made-up company, grouped into bands of 10 k€. Each brick is one employee. The curve through the tops of the columns gives the shape of the distribution: a peak around 40 to 50 k€ and a long tail toward high salaries.

Is the distribution balanced on both sides, or does it have a long tail stretching in one direction?

  • Symmetric: Skewness ≈ 0 (Mean ≈ Median ≈ Mode)
  • Right-skewed (positive): Skewness > 0 (The tail stretches to the right. Mode < Median < Mean)
  • Left-skewed (negative): Skewness < 0 (The tail stretches to the left. Mean < Median < Mode)

For a sample of nn values, the skewness coefficient is:

S=1n∑i=1n(xi−xˉ)3s3S = \frac{\frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^3}{s^3}

Where:

  • xix_i is each value in your sample.
  • xˉ\bar{x} is the arithmetic mean.
  • ss is the standard deviation.
  • nn is the sample size.

Reading the formula: why the power of 3?

This is where the whole statistical trick lies. Let’s look at the numerator, called the third moment: ∑(xi−xˉ)3\sum (x_i - \bar{x})^3.

  1. Distance from the mean: We first compute the distance between each value and the mean, (xi−xˉ)(x_i - \bar{x}).
  2. Cubing:
    • Unlike the variance (where we square, which makes everything positive), cubing keeps the sign.
    • If a value is far above the mean (on the right), the deviation is positive and its cube is strongly positive.
    • If a value is far below the mean (on the left), the deviation is negative and its cube is strongly negative.
  3. Amplifying the extremes: Cubing gives disproportionate “weight” to values far from the center. A single value far out in the tail will dominate the whole sum.

Normalizing by the standard deviation (s³)

We divide the result by s3s^3 to get a unitless (dimensionless) coefficient. This lets us compare the skewness of two different distributions, say salaries in euros and ages in years, because the result no longer depends on the scale of the data.

Interpreting the result

  • S≈0S \approx 0: Positive and negative cubed deviations roughly cancel out. The distribution is symmetric.
  • S>0S > 0: Large positive deviations win. The tail stretches to the right (positive skew).
  • S<0S < 0: Large negative deviations dominate. The tail stretches to the left (negative skew).
SkewnessS = 1.00
−3−2−10123Value (standardized)
Mean 0.00Median -0.16Mode -0.50

Mode < Median < Mean
The tail stretches to the right and pulls the mean along with it.

Figure 2: Move the slider to change the skewness coefficient S. Each curve is a standardized gamma distribution (mirrored when S < 0), so the mean stays at 0 and only the shape changes. The longer the tail, the further the mean moves away from the mode.

ML tip: Many machine learning models perform worse on skewed data. If your data is heavily skewed (like house prices, where a few luxury homes distort the mean), we often apply mathematical tricks to make it more symmetric before training.


Kurtosis

Skewness tells us about the symmetry of our data; kurtosis tells us about the tails of the distribution. More precisely: how likely are extreme values (outliers) compared with a normal distribution?

Definition

Kurtosis measures how heavy the extremities (tails) of a distribution are. It tells you whether your data has heavy tails (more outliers) or light tails (fewer outliers) compared with a normal distribution.

K=1n∑i=1n(xi−xˉ)4s4K = \frac{\frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^4}{s^4}

The fourth power amplifies extreme deviations, which makes kurtosis very sensitive to outliers.

  • Platykurtic (kurtosis < 3): flatter peak, thinner tails. Fewer extreme values than the normal distribution.
  • Mesokurtic (kurtosis = 3): normal distribution. The baseline for comparison.
  • Leptokurtic (kurtosis > 3): sharper peak, fatter tails. More extreme values than the normal distribution.
KurtosisK = 3.00
−3−2−10123Value (variance = 1)
Distribution shown (K = 3.00, excess +0.00)Normal distribution (reference)

Mesokurtic: this is the normal distribution, the reference (K = 3).
P(|X| > 2) = 4.55% · P(|X| > 3) = 0.27%
Normal distribution: 4.55% and 0.27%

Figure 3: Generalized normal distribution with variance 1. Every curve has the same variance; only the shape changes. The shaded areas beyond ±2 give the probability of an extreme value, to compare with the dotted normal curve.

If the value 3 puzzles you, it is simply the kurtosis of a standard normal distribution, used as a reference point to assess other distributions.

The density function of the standard normal distribution (μ=0,σ=1\mu = 0, \sigma = 1) is:

f(x)=12πe−x22f(x) = \frac{1}{\sqrt{2\pi}} e^{-\frac{x^2}{2}}

Kurtosis is defined as the expectation of the variable to the fourth power, E[X4]E[X^4]. Mathematically, this means computing the following integral:

I=12π∫−∞+∞x4e−x22 dxI = \frac{1}{\sqrt{2\pi}} \int_{-\infty}^{+\infty} x^4 e^{-\frac{x^2}{2}} \, dx

Since the function is even, we can restrict the study to [0,+∞[[0, +\infty[ and multiply the result by 2:

I=22π∫0+∞x4e−x22 dxI = \frac{2}{\sqrt{2\pi}} \int_{0}^{+\infty} x^4 e^{-\frac{x^2}{2}} \, dx

First Integration by Parts

We split the term x4e−x22x^4 e^{-\frac{x^2}{2}} into two parts to bring out the derivative of the exponent:

  • Let u=x3  ⟹  du=3x2 dxu = x^3 \implies du = 3x^2 \, dx
  • Let dv=xe−x22 dx  ⟹  v=−e−x22dv = x e^{-\frac{x^2}{2}} \, dx \implies v = -e^{-\frac{x^2}{2}}

The integration by parts formula (∫u dv=[uv]−∫v du\int u \, dv = [uv] - \int v \, du) gives:

∫0+∞x4e−x22 dx=[−x3e−x22]0+∞⏟=0−∫0+∞−3x2e−x22 dx\int_{0}^{+\infty} x^4 e^{-\frac{x^2}{2}} \, dx = \underbrace{\left[ -x^3 e^{-\frac{x^2}{2}} \right]_{0}^{+\infty}}_{= 0} - \int_{0}^{+\infty} -3x^2 e^{-\frac{x^2}{2}} \, dx

The first term vanishes because the exponential dominates the power at +∞+\infty. We are left with:

3∫0+∞x2e−x22 dx3 \int_{0}^{+\infty} x^2 e^{-\frac{x^2}{2}} \, dx

Second Integration by Parts

We repeat the process on the remaining integral ∫x2e−x22 dx\int x^2 e^{-\frac{x^2}{2}} \, dx:

  • Let u=x  ⟹  du=dxu = x \implies du = dx
  • Let dv=xe−x22 dx  ⟹  v=−e−x22dv = x e^{-\frac{x^2}{2}} \, dx \implies v = -e^{-\frac{x^2}{2}}

Integration by parts gives:

∫0+∞x2e−x22 dx=[−xe−x22]0+∞⏟=0−∫0+∞−e−x22 dx\int_{0}^{+\infty} x^2 e^{-\frac{x^2}{2}} \, dx = \underbrace{\left[ -x e^{-\frac{x^2}{2}} \right]_{0}^{+\infty}}_{= 0} - \int_{0}^{+\infty} -e^{-\frac{x^2}{2}} \, dx

We are then left with:

∫0+∞e−x22 dx\int_{0}^{+\infty} e^{-\frac{x^2}{2}} \, dx

This is the classic Gaussian integral. We know that ∫−∞+∞e−x22 dx=2π\int_{-\infty}^{+\infty} e^{-\frac{x^2}{2}} \, dx = \sqrt{2\pi}. By symmetry, over [0,+∞[[0, +\infty[ it equals 2π2\frac{\sqrt{2\pi}}{2}.

Putting it together

Let’s walk back up the chain of calculations:

  1. The second integration by parts gave us 2π2\frac{\sqrt{2\pi}}{2}.
  2. The first one multiplied that result by 33, giving 3×2π23 \times \frac{\sqrt{2\pi}}{2}.
  3. Finally, we must not forget the initial factor 22π\frac{2}{\sqrt{2\pi}} in front of the whole integral.

The final computation is therefore:

I=22π×(3×2π2)I = \frac{2}{\sqrt{2\pi}} \times \left( 3 \times \frac{\sqrt{2\pi}}{2} \right)

Simplifying by 22 and by 2π\sqrt{2\pi}, we get:

I=3\boxed{I = 3}

This shows that for a normal distribution, the ratio between the spread of the extremes (fourth moment) and the squared variance is exactly the constant 3.


Measures of Dispersion

We have studied the center (mean, median) and the shape (skewness). But one essential piece of the puzzle is still missing: dispersion.

Imagine two companies, A and B, with exactly the same average salary of $100,000.

  • At Company A, pay is perfectly equal: everyone earns $100,000.
  • At Company B, things are different: the CEO earns $460,000 and the four other employees earn $10,000 each.

The mean is identical ($100k), but the dispersion tells a completely different story. So we need a tool to measure how far, on average, the data points lie from the center.


Variance and Standard Deviation

These are the most widely used measures of dispersion in statistics and data science. They quantify how far each data point lies from the mean.

Variance (σ²)

The variance is the mean of the squared deviations from the mean. The higher the variance, the more spread out the data.

σ2=1n∑i=1n(xi−xˉ)2\sigma^2 = \frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^2

Standard Deviation (σ)

The standard deviation is the square root of the variance. Its main advantage is that it is expressed in the same unit as the original data (hours or dollars, for example), which makes it easier to interpret.

σ=σ2=1n∑i=1n(xi−xˉ)2\sigma = \sqrt{\sigma^2} = \sqrt{\frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^2}

Worked example: Light bulb lifetimes

You test 5 light bulbs with the following lifetimes: 980, 1000, 1010, 1020 and 1040 hours.

  1. Computing the mean (xˉ\bar{x}):
xˉ=980+1000+1010+1020+10405=1010 hours\bar{x} = \frac{980 + 1000 + 1010 + 1020 + 1040}{5} = 1010 \text{ hours}
  1. Computing the deviations from the mean (xi−xˉx_i - \bar{x}):

    • 980−1010=−30980 - 1010 = -30
    • 1000−1010=−101000 - 1010 = -10
    • 1010−1010=01010 - 1010 = 0
    • 1020−1010=+101020 - 1010 = +10
    • 1040−1010=+301040 - 1010 = +30
  2. Squaring the deviations:

(−30)2=900∣(−10)2=100∣(0)2=0∣(10)2=100∣(30)2=900(-30)^2 = 900 \quad | \quad (-10)^2 = 100 \quad | \quad (0)^2 = 0 \quad | \quad (10)^2 = 100 \quad | \quad (30)^2 = 900
  1. Computing the variance (σ2\sigma^2):
σ2=900+100+0+100+9005=400 hours2\sigma^2 = \frac{900 + 100 + 0 + 100 + 900}{5} = 400 \text{ hours}^2
  1. Computing the standard deviation (σ\sigma):
σ=400=20 hours\sigma = \sqrt{400} = 20 \text{ hours}

Why square the deviations?

There are two fundamental reasons for this mathematical choice:

  1. Neutralizing signs: If we simply added up the raw deviations, positive and negative values would always cancel out. Squaring makes every distance positive.
  2. Penalizing extreme values: Squaring gives disproportionate weight to large deviations. This makes the variance and the standard deviation very sensitive to outliers.

Range and Interquartile Range (IQR)

Range

This is the simplest measure of dispersion. It is the difference between the maximum and the minimum value of a dataset.

Range=Max−Min\text{Range} = \text{Max} - \text{Min}
  • Weakness: It is extremely sensitive to outliers.

Interquartile Range (IQR)

The interquartile range focuses on the spread of the middle 50% of the data. It is the difference between the third quartile (Q3Q_3) and the first quartile (Q1Q_1).

IQR=Q3−Q1\text{IQR} = Q_3 - Q_1
  • Components:

    • Q1Q_1 (25th percentile): The value below which 25% of the data lies.
    • Q3Q_3 (75th percentile): The value below which 75% of the data lies.
  • Strength: It is robust to outliers.


Detecting Outliers

A widely accepted rule flags as an outlier any data point that falls outside the “fences” computed from the interquartile range (IQR).

The 1.5 × IQR Rule

We define two critical thresholds to identify these anomalies:

  • Lower fence: Any value below Q1−1.5×IQRQ_1 - 1.5 \times \text{IQR}.
  • Upper fence: Any value above Q3+1.5×IQRQ_3 + 1.5 \times \text{IQR}.

This is the rule used to draw the whiskers of box plots.

Q1Q2Q3whiskerwhiskermeanoutliers0102030405060708090Response time (ms)
Sorted measurements12118221323424526627728829930103111321234133514371539164217451871198420
The complete box plot
1/9

At the top, each of the 20 measurements is a dot. Below, the box plot summarizes them: the box runs from Q1 to Q3, the line inside it is the median, the whiskers cover the ordinary measurements and the isolated circles are the outliers. The next steps compute each element.

Figure 4: Box plot of 20 server response times (made-up data), built step by step. The quartiles follow the definition of empirical quantiles given above; software that interpolates (NumPy, Excel) may give slightly different values.

Origin and Rationale: Why 1.5?

The 1.5 coefficient was popularized by the statistician John Tukey. It is not an absolute mathematical law but a heuristic (a rule of thumb) based on a statistical compromise tied to the normal distribution. The idea was to strike a balance between flagging outliers too eagerly and letting them slip through.


Formula Reference

MetricFormulaWhen to use it
Meanxˉ=1n∑xi\bar{x} = \frac{1}{n}\sum x_iData is symmetric, with no extreme outliers
MedianMiddle value of the sorted dataData is skewed or contains outliers
ModeMost frequent valueCategorical data or finding peaks
Varianceσ2=1n∑(xi−xˉ)2\sigma^2 = \frac{1}{n}\sum(x_i - \bar{x})^2Measuring spread (squared units)
Standard deviationσ=σ2\sigma = \sqrt{\sigma^2}Measuring spread (original units)
SkewnessS=1n∑(xi−xˉ)3s3S = \frac{\frac{1}{n}\sum(x_i-\bar{x})^3}{s^3}Checking the symmetry of the distribution
KurtosisK=1n∑(xi−xˉ)4s4K = \frac{\frac{1}{n}\sum(x_i-\bar{x})^4}{s^4}Checking how heavy the tails are
RangeMax−Min\text{Max} - \text{Min}Quick estimate of the spread
IQRQ3−Q1Q_3 - Q_1Robust spread, outlier detection

Weekly Notes

Every Sunday, I share what I've been learning: papers, ideas, experiments, and questions that stayed with me.

You can unsubscribe at any time with a single click.

0 Likes • 0 Comments

Discussion about this post0

Join the discussion

A secure sign-in link will be sent to your email address.

Loading discussion...