A penalty has two outcomes and nothing in between. That simple idea, the Bernoulli trial, is the building block of expected goals and of most football statistics.
One penalty is a Bernoulli trial. Five penalties, each with the same chance, is the Binomial distribution, and it shows why missing one in five is normal, not a slump.
A team averages two goals a game. How likely is it to score exactly three? The Poisson distribution turns an average into a probability for every score, and it sits underneath most football prediction models.
A striker scores with one shot in five. How many shots until his first goal? The Geometric distribution answers it, and shows why a three-match drought is often just bad luck.
A match has three possible results, not two. The Multinomial distribution handles any number of outcomes, and shows why a team's "expected" record over ten games almost never happens exactly.
Ranking players from fastest to slowest tells you the order, not how unusual anyone is. The Normal distribution, and its standard deviation, measures how far a player stands out from the rest.
The Uniform distribution says every outcome is equally likely. It's the fairest-sounding model in statistics, and testing it against 28,016 goals shows football isn't that fair.
A player's performance can be written as a list of numbers in a fixed order. That list is a vector, and it's the first step to comparing players, finding replacements and feeding football into machine learning.
One player's performance is a vector. Stack several players together and you have a matrix, the grid that almost all of data science starts from. Multiply it by a set of weights and every player gets a rating.
Two players' stats are two vectors, and the distance between them measures how alike they are. It's how recruitment teams shortlist replacements, and it only works once every stat is put on the same scale.
A run from one spot on the pitch to another is a vector, how far forward and how far across. From that pair of numbers you get the length of the run, its direction, its speed and whether it was heading for goal.
Every pass has a start and an end, so every pass is a vector. That turns "he never passes forward" from an opinion into a number, and shows why a player's average pass can hide half of what he does.
A new signing scores his first four penalties. Bayesian thinking combines what you believed before with what you've just seen, and shows why a perfect start should nudge your opinion, not replace it.
BeginnerPart 1 of Bayesian Thinking Through Football
A half-time lead is evidence, and what it means depends on who's leading. Bayes' theorem combines the pre-match odds with the half-time score, and 5,419 SPFL matches show it working.
BeginnerPart 3 of Bayesian Thinking Through Football
Every machine learning model starts with two decisions, what to predict and what to tell the model. The first is the target, the second the features, and real SPFL results show why the features matter as much as the model.
BeginnerPart 1 of Machine Learning Through Football
A model that has seen the answers will always look brilliant. Holding matches back to test it, and choosing which ones, is how you find out whether it can predict a match it hasn't seen. Real SPFL seasons show the difference.
BeginnerPart 2 of Machine Learning Through Football
A model can learn its training matches too well, rules and all, and then fall apart on matches it hasn't seen. Real SPFL seasons show how it happens, how to spot it, and what to do about it.
BeginnerPart 3 of Machine Learning Through Football
A model can fail by being too simple to see the patterns or too complicated to ignore the noise. Real SPFL seasons, and the bookmakers, show what each looks like and where the sweet spot sits.
BeginnerPart 4 of Machine Learning Through Football
A model can be too rigid to learn, or so sensitive that one result rewrites everything it believes. Those are bias and variance, and twenty-five years of SPFL results show what each costs.
BeginnerPart 5 of Machine Learning Through Football
A single test season can flatter a model or bury it. Cross-validation tests it again and again on different slices of the data; for football, that means walking forward through the seasons. 23 SPFL seasons show why it matters.
BeginnerPart 6 of Machine Learning Through Football
Accuracy counts how many results a model called right, and hides almost everything else. A confusion matrix shows where the mistakes are, probability scores show how confident it was, and five SPFL seasons show why draws break accuracy.
BeginnerPart 7 of Machine Learning Through Football
Every time you say a winger shot away from a defender, you're talking about rate of change. That's the heart of calculus, and a few seconds of sprinting are enough to see the derivative at work.
Still accelerating, at full speed, lost a yard of pace. Each is about what's happening at one moment, which is what the derivative measures, and it gives coaches something to train, measure and improve.
Two players can hit the same top speed and look identical for a second, then one pulls clear. The difference is how quickly their acceleration fades, and that is the second derivative.
Speed tells you how fast; integration tells you how far. It's the area under the speed graph, and it's how tracking data turns thousands of speed readings into distance covered and high-speed running.
One run into space, pulled apart and put back together. Position, speed and acceleration are one chain linked by derivatives and integrals, and it's how tracking data turns dots on a pitch into what coaches want to know.
Five players on a shortlist, £10m to spend. Chase goals and you sign the striker. Chase points and three cheaper players beat him. Add one rule and the answer changes again. The three parts of every optimisation problem, with a transfer window.
BeginnerPart 1 of Decision Science Through Football