Supervised learning. Each learns from past matches with known results, using three numbers known before kick-off, and is tested on five seasons it never saw (log loss, lower is better).
Logistic regression
Football question: Home win, draw or away win, and how likely is each?
How it works: Adds up each feature times a weight, then turns the totals into three probabilities that sum to 100%.
Good at: Simple, fast, readable weights. Scored 0.950, about 90% of the way from knowing nothing to the bookmakers, level with KNN as the best learner here.
Watch out for: Each feature counts in a straight line, and it never picks a draw as the most likely result.
Don't use it when: Effects bend or only kick in past a threshold, unless you build that into the features.
Decision tree
Football question: Which few questions sort matches into home wins, draws and away wins?
How it works: Asks yes/no questions, each chosen because it splits the matches best, then predicts the results in each final group.
Good at: Reads like a pundit's reasoning. Two questions deep it scored 0.954, close to logistic regression.
Watch out for: Grown deep it memorises. Twelve questions deep, 67% right on training matches, 47% on new ones.
Don't use it when: You need smooth forecasts. Two almost identical matches either side of a cut get very different ones.
Random forest
Football question: Can a hundred jumpy trees make one steady forecast?
How it works: Grows many trees, each on a resampled set of matches with a random choice of question, and averages their forecasts.
Good at: Fixes most of a single tree's wild guesses. A forest of 100 scored 0.953.
Watch out for: Hard to read, and scores shift slightly with the random seed. Still just behind logistic regression.
Don't use it when: You only have a few features. Its random choice of question is then more handicap than help.
Gradient boosting
Football question: Can each new model fix what the last one got wrong?
How it works: Starts with a rough guess, then adds hundreds of small trees, each fitted to the remaining mistakes.
Good at: One of the strongest methods on tables with many columns. Here it drew level with logistic regression at 0.951.
Watch out for: Too many rounds overfit, so choose the number on held-back seasons, never the test.
Don't use it when: The features are the limit. The four learners of parts 8 to 11 finished within 0.004 of each other.
K-nearest neighbours (KNN)
Football question: Which past matches were most like this one, and how did they go?
How it works: Measures the distance from this match to every past match on the same gaps, and counts the results of the k nearest.
Good at: Learns nothing, yet drew level with logistic regression at 0.949, and gets the evenly matched game nearly right.
Watch out for: Too few neighbours is mostly luck. Five did worse than knowing nothing; 300 was best.
Don't use it when: Some features are useless. Every feature gets the same say, and it can't learn to ignore one.
Naive Bayes
Football question: Starting from how often each result happens, how should each clue shift the odds?
How it works: Multiplies the base rates by how typical each clue is of home wins, draws and away wins, one clue at a time, as if each were separate news.
Good at: Quick, simple and easy to follow clue by clue. With one combined strength clue it scored 0.951, level with logistic regression.
Watch out for: Overlapping clues get counted twice. With three gaps it said 92% for home wins that happened 80% of the time, and scored 1.020.
Don't use it when: The features repeat each other. Merge them into one first, or use a model that fits the weights together.