Football stats overlap. Principal component analysis finds the few directions that carry most of the information. On 26 seasons of Scottish Premiership data, one component built from eight stats explains 56% of the variation and tracks points per game better than any single stat.
Every club's home record says something about its ground, and a lot about luck. Partial pooling lets the data decide how far to trust each club's own figure. Over 26 Scottish Premiership seasons the clubs differ far less than the raw table suggests, and Hearts still come out on top.
AdvancedPart 6 of Bayesian Thinking Through Football
Four machine learning models stalled at the same score on three numbers per match. Rate every team's attack and defence from the goals they score and let in, refit before every matchday, and a classic football model finally gets past them.
AdvancedPart 12 of Machine Learning Through Football
Elo has dials to set, from how far one result moves a rating to how much home advantage is worth. Tuned on held-back seasons it draws level with Dixon-Coles on the test seasons. The dial that mattered most changed between eras, and no honest tuning could have seen it coming.
AdvancedPart 13 of Machine Learning Through Football
A forecast can lose by being wrong about its own confidence, or by knowing less. Splitting the Brier score into calibration and resolution shows which. On five SPFL test seasons Elo is as well calibrated as the bookmakers, and loses only because they know more.
AdvancedPart 14 of Machine Learning Through Football