Skip to content

Learn

Statistics, maths and machine learning, each one introduced through a football question.

Every beginner piece, series by series, in part order.

Goal or miss? The Bernoulli distribution

A penalty has two outcomes and nothing in between. That simple idea, the Bernoulli trial, is the building block of expected goals and of most football statistics.

Beginner Part 1 of Statistics Through Football

Five penalties, how many go in? The Binomial distribution

One penalty is a Bernoulli trial. Five penalties, each with the same chance, is the Binomial distribution, and it shows why missing one in five is normal, not a slump.

Beginner Part 2 of Statistics Through Football

How many goals will we score? The Poisson distribution

A team averages two goals a game. How likely is it to score exactly three? The Poisson distribution turns an average into a probability for every score, and it sits underneath most football prediction models.

Beginner Part 3 of Statistics Through Football

How many shots until he scores? The Geometric distribution

A striker scores with one shot in five. How many shots until his first goal? The Geometric distribution answers it, and shows why a three-match drought is often just bad luck.

Beginner Part 5 of Statistics Through Football

Win, draw or lose? The Multinomial distribution

A match has three possible results, not two. The Multinomial distribution handles any number of outcomes, and shows why a team's "expected" record over ten games almost never happens exactly.

Beginner Part 7 of Statistics Through Football

How fast is fast? The Normal distribution

Ranking players from fastest to slowest tells you the order, not how unusual anyone is. The Normal distribution, and its standard deviation, measures how far a player stands out from the rest.

Beginner Part 10 of Statistics Through Football

A midfielder's match in five numbers. Vectors

A player's performance can be written as a list of numbers in a fixed order. That list is a vector, and it's the first step to comparing players, finding replacements and feeding football into machine learning.

Beginner Part 1 of Linear Algebra Through Football

Four players, four numbers each. The whole team becomes a matrix

One player's performance is a vector. Stack several players together and you have a matrix, the grid that almost all of data science starts from. Multiply it by a set of weights and every player gets a rating.

Beginner Part 2 of Linear Algebra Through Football

Who plays most like him? Measuring player similarity with distance

Two players' stats are two vectors, and the distance between them measures how alike they are. It's how recruitment teams shortlist replacements, and it only works once every stat is put on the same scale.

Beginner Part 3 of Linear Algebra Through Football

Where did he run? Describing player movement with vectors

A run from one spot on the pitch to another is a vector, how far forward and how far across. From that pair of numbers you get the length of the run, its direction, its speed and whether it was heading for goal.

Beginner Part 5 of Linear Algebra Through Football

Forward or sideways? Passing as a vector

Every pass has a start and an end, so every pass is a vector. That turns "he never passes forward" from an opinion into a number, and shows why a player's average pass can hide half of what he does.

Beginner Part 6 of Linear Algebra Through Football

Four from four. How good is he really? Bayesian thinking

A new signing scores his first four penalties. Bayesian thinking combines what you believed before with what you've just seen, and shows why a perfect start should nudge your opinion, not replace it.

Beginner Part 1 of Bayesian Thinking Through Football

1–0 up at half-time. Will they hold on? Bayes' theorem

A half-time lead is evidence, and what it means depends on who's leading. Bayes' theorem combines the pre-match odds with the half-time score, and 5,419 SPFL matches show it working.

Beginner Part 3 of Bayesian Thinking Through Football

What are we trying to predict? Features and targets

Every machine learning model starts with two decisions, what to predict and what to tell the model. The first is the target, the second the features, and real SPFL results show why the features matter as much as the model.

Beginner Part 1 of Machine Learning Through Football

Has the model learned, or just memorised? Training data and test data

A model that has seen the answers will always look brilliant. Holding matches back to test it, and choosing which ones, is how you find out whether it can predict a match it hasn't seen. Real SPFL seasons show the difference.

Beginner Part 2 of Machine Learning Through Football

Brilliant in training, gone on matchday. Overfitting

A model can learn its training matches too well, rules and all, and then fall apart on matches it hasn't seen. Real SPFL seasons show how it happens, how to spot it, and what to do about it.

Beginner Part 3 of Machine Learning Through Football

Too simple or too clever? Underfitting vs overfitting

A model can fail by being too simple to see the patterns or too complicated to ignore the noise. Real SPFL seasons, and the bookmakers, show what each looks like and where the sweet spot sits.

Beginner Part 4 of Machine Learning Through Football

Stubborn or jumpy? Bias and variance

A model can be too rigid to learn, or so sensitive that one result rewrites everything it believes. Those are bias and variance, and twenty-five years of SPFL results show what each costs.

Beginner Part 5 of Machine Learning Through Football

One good season or a good model? Cross-validation

A single test season can flatter a model or bury it. Cross-validation tests it again and again on different slices of the data; for football, that means walking forward through the seasons. 23 SPFL seasons show why it matters.

Beginner Part 6 of Machine Learning Through Football

Is accuracy the right score? Evaluating a model

Accuracy counts how many results a model called right, and hides almost everything else. A confusion matrix shows where the mistakes are, probability scores show how confident it was, and five SPFL seasons show why draws break accuracy.

Beginner Part 7 of Machine Learning Through Football

Still accelerating, or at full speed? The derivative

Still accelerating, at full speed, lost a yard of pace. Each is about what's happening at one moment, which is what the derivative measures, and it gives coaches something to train, measure and improve.

Beginner Part 2 of Calculus Through Football

Is his acceleration fading? The second derivative

Two players can hit the same top speed and look identical for a second, then one pulls clear. The difference is how quickly their acceleration fades, and that is the second derivative.

Beginner Part 3 of Calculus Through Football

We know how fast he was running. How far did he get? Integration

Speed tells you how fast; integration tells you how far. It's the area under the speed graph, and it's how tracking data turns thousands of speed readings into distance covered and high-speed running.

Beginner Part 4 of Calculus Through Football

Who's the best signing? It depends what you're optimising

Five players on a shortlist, £10m to spend. Chase goals and you sign the striker. Chase points and three cheaper players beat him. Add one rule and the answer changes again. The three parts of every optimisation problem, with a transfer window.

Beginner Part 1 of Decision Science Through Football