Skip to content

Cheat sheet

Machine learning methods, football edition. What each one answers, where it shines, where it breaks, and where to try it. Every row comes from what our own articles found on real matches, not from a textbook.

Scores are log loss on five test seasons, 990 matches: lower is better, and the bookmakers score 0.932.

Predicting results

Supervised learning. Each learns from past matches with known results, using three numbers known before kick-off, and is tested on five seasons it never saw (log loss, lower is better).

Logistic regression

Football question: Home win, draw or away win, and how likely is each?

How it works: Adds up each feature times a weight, then turns the totals into three probabilities that sum to 100%.

Good at: Simple, fast, readable weights. Scored 0.950, about 90% of the way from knowing nothing to the bookmakers, level with KNN as the best learner here.

Watch out for: Each feature counts in a straight line, and it never picks a draw as the most likely result.

Don't use it when: Effects bend or only kick in past a threshold, unless you build that into the features.

Decision tree

Football question: Which few questions sort matches into home wins, draws and away wins?

How it works: Asks yes/no questions, each chosen because it splits the matches best, then predicts the results in each final group.

Good at: Reads like a pundit's reasoning. Two questions deep it scored 0.954, close to logistic regression.

Watch out for: Grown deep it memorises. Twelve questions deep, 67% right on training matches, 47% on new ones.

Don't use it when: You need smooth forecasts. Two almost identical matches either side of a cut get very different ones.

Random forest

Football question: Can a hundred jumpy trees make one steady forecast?

How it works: Grows many trees, each on a resampled set of matches with a random choice of question, and averages their forecasts.

Good at: Fixes most of a single tree's wild guesses. A forest of 100 scored 0.953.

Watch out for: Hard to read, and scores shift slightly with the random seed. Still just behind logistic regression.

Don't use it when: You only have a few features. Its random choice of question is then more handicap than help.

Gradient boosting

Football question: Can each new model fix what the last one got wrong?

How it works: Starts with a rough guess, then adds hundreds of small trees, each fitted to the remaining mistakes.

Good at: One of the strongest methods on tables with many columns. Here it drew level with logistic regression at 0.951.

Watch out for: Too many rounds overfit, so choose the number on held-back seasons, never the test.

Don't use it when: The features are the limit. The four learners of parts 8 to 11 finished within 0.004 of each other.

K-nearest neighbours (KNN)

Football question: Which past matches were most like this one, and how did they go?

How it works: Measures the distance from this match to every past match on the same gaps, and counts the results of the k nearest.

Good at: Learns nothing, yet drew level with logistic regression at 0.949, and gets the evenly matched game nearly right.

Watch out for: Too few neighbours is mostly luck. Five did worse than knowing nothing; 300 was best.

Don't use it when: Some features are useless. Every feature gets the same say, and it can't learn to ignore one.

Naive Bayes

Football question: Starting from how often each result happens, how should each clue shift the odds?

How it works: Multiplies the base rates by how typical each clue is of home wins, draws and away wins, one clue at a time, as if each were separate news.

Good at: Quick, simple and easy to follow clue by clue. With one combined strength clue it scored 0.951, level with logistic regression.

Watch out for: Overlapping clues get counted twice. With three gaps it said 92% for home wins that happened 80% of the time, and scored 1.020.

Don't use it when: The features repeat each other. Merge them into one first, or use a model that fits the weights together.

Predicting numbers

Supervised learning for an amount rather than a result, such as goals over a season. Judged on how far it misses on seasons it never saw.

Linear regression

Football question: How many shots is a goal worth?

How it works: Draws the straight line through the data that makes the squared misses smallest; its slope is what one more of something is worth.

Good at: Every weight reads as a number. One goal every 3.1 shots on target, and several features at once, each with the others held the same.

Watch out for: The world can change under it. Shots alone fell from 65% explained to 30% on the test seasons after recorded shots jumped.

Don't use it when: The relationship bends, or what you measure is recorded differently from the seasons it learned on.

Football's own models

Built for football scores rather than borrowed from general machine learning. Each works from goals or results, and the two rating models beat every learner above.

Poisson goals

Football question: A team averages two goals a game. How likely is exactly three?

How it works: Turns an average scoring rate into a probability for every number of goals, and so for every scoreline.

Good at: Sits underneath most football prediction models, and needs only one number per team.

Watch out for: Assumes one steady rate for 90 minutes and two independent teams, so it tends to under-rate 0–0 and 1–1.

Don't use it when: You can't estimate the scoring rate well. Poisson is only as good as the average you feed it.

Dixon-Coles ratings

Football question: How good is each team's attack, and each team's defence?

How it works: Gives every team an attack and a defence rating plus home advantage, fitted to recent goals and refitted every matchday.

Good at: Scored 0.945, the first model in the series past logistic regression's 0.950.

Watch out for: Goals are noisy, and promoted sides arrive with little history to rate them on.

Don't use it when: Team news matters most. Injuries, suspensions and a new manager are invisible to it.

Elo ratings

Football question: How far should one result move a team's rating?

How it works: Each side has one rating; the gap gives a win expectancy, and every result moves both ratings by how surprising it was.

Good at: Simple to keep up to date. Tuned honestly on held-back seasons it matched Dixon-Coles at 0.945.

Watch out for: Its settings can go stale. Home advantage rose in the test seasons, which no honest tuning could have seen coming.

Don't use it when: Grounds differ a lot. One home advantage for everyone is a simplification.

Finding patterns

Unsupervised learning. No results to predict; the method looks for structure in the numbers themselves.

Principal component analysis (PCA)

Football question: Eight team stats overlap. Is there one story underneath them?

How it works: Finds the few weighted mixes of the stats that carry most of the differences between teams.

Good at: One component built from eight stats explains 56% of the variation, and tracks points per game better than any single stat.

Watch out for: The names are ours. PCA gives weights, not labels, and later components may mean nothing in football terms.

Don't use it when: The biggest source of variation isn't the one you care about. Fouls vary a lot but don't decide matches.

Not on this site, and why

Famous methods we haven't used, with the honest reason.

Neural networks (MLP, CNN, RNN, transformers)
They shine on images, video, text and tracking data, with far more examples than 26 seasons of results. On three numbers per match the learners above already hit the ceiling.
Support vector machines
On three features they draw much the same kind of boundary as logistic regression, and don't give probabilities without extra work.
DBSCAN
Finds groups of any shape by how densely points pack together. It needs plenty of points to tell dense from sparse, and a few hundred team-seasons is little to go on.