Four ways to play. K-means and team styles
Give a computer eight stats for 310 Scottish team-seasons and no labels, and K-means sorts them into four styles, from dominant to outshot. The outshot sides who foul a lot take no more points than the ones who don't.
Intermediate Part 18 of Machine Learning Through Football
New to the notation? The symbols explained
Contents
The football question
Pundits sort teams into types without thinking: the dominant side, the ones who sit in and get battered, the ones who make it a scrap. If you gave a computer every team's numbers and no labels, would it find the same types? And does the scrap actually work?
Finding groups without being told what they are is clustering, and K-means is the classic way to do it. Every model so far in this series was told the right answers, the results, and learned to predict them. This one gets no answers at all. That's unsupervised learning.
The data
The same eight numbers as PCA, for the same 310 team-seasons, every Scottish Premiership side from 2000/01 to 2025/26 with the stats recorded in at least 30 matches: shots, shots on target and corners, both for and against, plus fouls and yellow cards, all per game.
Each stat is then measured in standard deviations, so that shots and yellow cards count on the same scale, the step from measuring player similarity. But measured against what? That turns out to matter.
The concept
Choose a number of groups, k. Then K-means looks for the k group centres that make every team-season as close as possible to its own group's centre:
$$\text{spread} = \sum d(\text{team}, \text{centre})^2$$
In plain football
- d is the distance between a team-season's eight numbers and a centre's, the same straight-line distance as k-nearest neighbours.
- A centre is the average style of the team-seasons in its group: a made-up "typical" side for that style.
- Σ means add up over every team-season, and spread is that total: how far each team-season sits from its own group's centre, squared. Small spread means tight, clear groups.
It gets there in a loop. Start with k team-seasons picked at random as centres. Put every team-season in the group of its nearest centre. Move each centre to the average of its group. Repeat until nothing moves. Different random starts can end in different places, so it's run 20 times and the tightest result kept.
The era trap
The obvious way to standardise is against all 310 team-seasons together. Do that and ask for four groups, and one group contains no team-seasons from 2020/21 onwards while another is 52% from those seasons, against 23% overall.
That's not style. It's the change in how shots were recorded that broke a line in linear regression: from 2020/21 every side was credited with about two more shots a game. Measured against all seasons together, recent sides look like a different kind of team.
So each stat is measured against its own season's league instead: how many more shots than the average side that season. Then every group's share of recent seasons sits between 20% and 27%. The groups are about how teams played, not when.
How many groups?
K-means has to be told k. The spread always falls as k rises, since more groups means everyone is closer to some centre, so the question is where adding groups stops helping much:
| Groups, k | Spread inside the groups |
|---|---|
| 1 | 2480 |
| 2 | 1283 |
| 3 | 1007 |
| 4 | 845 |
| 5 | 765 |
| 6 | 707 |
The biggest drops come early and there's no sharp elbow, so k is a judgement call, and it's worth saying so. Four gives groups that are clear and different. And they're stable: fit K-means on the first 13 seasons alone, then on the last 13 alone, and the same four groups appear both times.
Four ways to play
| Group | Team-seasons | Points a game |
|---|---|---|
| Dominant | 50 | 2.29 |
| Middle | 102 | 1.35 |
| Outshot, clean | 79 | 1.11 |
| Outshot, physical | 79 | 1.10 |
- Dominant: 15.0 shots a game against 7.5, 7.0 corners against 3.8, and the fewest fouls. It's Celtic in all 26 seasons and Rangers in 20, plus four others: Aberdeen in 2014/15 and 2016/17, Hibernian in 2017/18 and Hearts in 2025/26, the season they led the league at halfway.
- Middle: close to level on everything, about 10 shots for and 10 against. Hibs, Aberdeen and Hearts are here most often.
- Outshot, clean: 9.2 shots a game against 11.9, and only 11.4 fouls. Dundee most often.
- Outshot, physical: outshot just as badly, 9.3 against 12.0, but 13.4 fouls and 2.0 yellow cards a game against the clean group's 11.4 and 1.5. Ross County most often.
Nobody told K-means about dominant sides or scraps. It found them in the numbers.
Does getting stuck in work?
The two outshot groups are a natural experiment: equally outshot, very different in how they react. The clean group takes 1.11 points a game, the physical group 1.10. The difference is +0.01, with a luck margin of ±0.08. No sign at all that fouling your way through it earns points.
Across all 310 team-seasons, physical play does go with fewer points: a correlation of −0.42. But that's because the dominant sides rarely need to foul, not because fouling costs points. Among the outshot teams only, the correlation is −0.01. Compare like with like, and the link vanishes: the same lesson as the era trap, and as beating the line in linear regression. Holding dominance fixed directly tells the same story, in do football stats mislead?
A family tree of clubs
K-means needs k up front. Hierarchical clustering doesn't: it starts with every club on its own and keeps joining the two closest groups until there's one. Take the 18 clubs with five or more Premiership seasons, give each its average style across those seasons, and measure how far apart two groups are as the average distance between their members:
Hearts and Hibernian are the second most alike pair in the league, joined at 0.62, behind only Hamilton and Livingston at 0.55. Whatever the derby feels like, over the seasons the Edinburgh clubs have played remarkably alike. They then join Aberdeen and Inverness. Celtic and Rangers aren't especially alike each other, joined at 1.47, but they're so far from everyone else that they join the rest of the league only at 5.51, three times further than any other join.
Why it matters
- Finding structure without labels. Customer types, player roles, styles of play: most data comes without a right answer to learn from, and clustering is how you start.
- Check what the groups really are. The first attempt grouped teams by era, because of how the data was recorded. Always ask what else could explain your groups.
- The groups make comparisons fair. "Teams that foul a lot take fewer points" is true and misleading. Comparing within a group, equally outshot sides, gives the real answer.
- k is a choice. There's no right number of styles in nature. Pick one that's clear, stable and useful, and say it was a choice.
Limitations
- Eight stats aren't a style. No possession, pressing, passing or chance quality: this data doesn't have them. With them, the groups would be finer.
- Season averages hide change. A side that changed manager in January gets one average for two styles.
- K-means likes round groups. It assumes groups are roughly ball-shaped and similar in size; long, thin or very uneven groups suit other methods.
- Fouling's cost is about points. It may matter in other ways, such as suspensions, that points a game can't show.
Try it yourself
The Team Style Map plots every team-season by dominance and physical play, coloured by group. Pick your club to light up its seasons, and tap one to see the five team-seasons most like it.
Or by hand: name the three Premiership sides you think play most like your club. Then find your club in the family tree above. Did you pick its neighbours?
Reproduce the analysis
The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It takes about ten seconds, most of it the 20 restarts:
Show the Python108 lines, ready to copy and run.
import csv
import random
from collections import Counter, defaultdict
from math import sqrt
from statistics import correlation, mean, pstdev
STATS = ["shots", "shots against", "on target", "on target against", "corners", "corners against", "fouls", "yellows"]
NEED = ("FTHG", "FTAG", "HS", "AS", "HST", "AST", "HC", "AC", "HF", "AF", "HY", "AY")
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]
# the eight per-game averages of Linear Algebra #9 for every team-season with the stats in 30+ matches, plus points
totals = defaultdict(lambda: [0] * 10)
for s in names:
with open(f"SC0_{s}.csv", encoding="latin-1") as f:
for r in csv.DictReader(f):
if not all(r.get(c) for c in NEED):
continue
v = {c: int(r[c]) for c in NEED}
for team, us, them in ((r["HomeTeam"], "H", "A"), (r["AwayTeam"], "A", "H")):
gf, ga = v[f"FT{us}G"], v[f"FT{them}G"]
add = [v[us + "S"], v[them + "S"], v[us + "ST"], v[them + "ST"], v[us + "C"], v[them + "C"],
v[us + "F"], v[us + "Y"], 3 if gf > ga else 1 if gf == ga else 0, 1]
totals[s, team] = [a + b for a, b in zip(totals[s, team], add)]
rows = [{"season": s, "team": team, "x": [t[i] / t[9] for i in range(8)], "ppg": t[8] / t[9]}
for (s, team), t in totals.items() if t[9] >= 30]
print(f"{len(rows)} team-seasons; 2020/21 on: {sum(r['season'] >= '2021' for r in rows)}")
def standardise(key, group_of): # each stat in standard deviations from the average of its group
groups = defaultdict(list)
for r in rows:
groups[group_of(r)].append(r)
for g in groups.values():
m = [mean(r["x"][j] for r in g) for j in range(8)]
sd = [pstdev([r["x"][j] for r in g]) for j in range(8)]
for r in g:
r[key] = [(r["x"][j] - m[j]) / sd[j] for j in range(8)]
standardise("all", lambda r: 0) # against all 310 together
standardise("z", lambda r: r["season"]) # against that season's league
def dist2(p, q):
return sum((a - b) ** 2 for a, b in zip(p, q))
def kmeans(points, k, seed): # start from k random team-seasons, then assign and re-centre until nothing moves
centres = [p[:] for p in random.Random(seed).sample(points, k)]
while True:
labels = [min(range(k), key=lambda c: dist2(p, centres[c])) for p in points]
new = [[mean(p[j] for p, l in zip(points, labels) if l == c) for j in range(8)] if c in labels else centres[c]
for c in range(k)]
if new == centres:
return labels, centres, sum(dist2(p, centres[l]) for p, l in zip(points, labels))
centres = new
def best_kmeans(points, k): # the tightest of 20 different starts
return min((kmeans(points, k, seed) for seed in range(20)), key=lambda run: run[2])
# 1. the era trap: four groups from the raw numbers, then from numbers measured against each season
for key, label in (("all", "against all seasons together"), ("z", "against each season")):
labels, _, _ = best_kmeans([r[key] for r in rows], 4)
shares = sorted(sum(r["season"] >= "2021" for r, l in zip(rows, labels) if l == c) / labels.count(c) for c in range(4))
print(f"{label}: share of each group from 2020/21 on: " + ", ".join(f"{s:.0%}" for s in shares))
# 2. how many groups? the spread left inside the groups for each k
points = [r["z"] for r in rows]
for k in range(1, 8):
print(f"k = {k}: spread inside the groups {best_kmeans(points, k)[2]:.0f}")
# 3. four groups, named by rule: most points dominant, next middle; of the outshot two, fewer fouls is clean
labels, centres, _ = best_kmeans(points, 4)
by_points = sorted(range(4), key=lambda c: -mean(r["ppg"] for r, l in zip(rows, labels) if l == c))
dominant, middle = by_points[:2]
clean, physical = sorted(by_points[2:], key=lambda c: centres[c][6])
NAMES = {dominant: "dominant", middle: "middle", clean: "outshot, clean", physical: "outshot, physical"}
dominance = lambda z: (z[0] + z[2] + z[4] - z[1] - z[3] - z[5]) / 6 # shots, on target, corners: for minus against
rough = lambda z: (z[6] + z[7]) / 2 # fouls and yellows
for c in (dominant, middle, clean, physical):
members = [r for r, l in zip(rows, labels) if l == c]
per_game = ", ".join(f"{s} {mean(r['x'][j] for r in members):.1f}" for j, s in enumerate(STATS))
print(f"{NAMES[c]}: {len(members)} team-seasons, {mean(r['ppg'] for r in members):.2f} points a game; "
f"dominance {dominance(centres[c]):+.2f}, physical {rough(centres[c]):+.2f}\n per game: {per_game}\n "
+ ", ".join(f"{t} {n}" for t, n in Counter(r["team"] for r in members).most_common(5)))
print("dominant, not Celtic or Rangers:", ", ".join(f"{r['team']} 20{r['season'][:2]}/{r['season'][2:]}"
for r, l in zip(rows, labels) if l == dominant and r["team"] not in ("Celtic", "Rangers")))
# 4. does fouling buy points? the two outshot groups, then the correlation with and without the dominant sides
a, b = ([r["ppg"] for r, l in zip(rows, labels) if l == c] for c in (clean, physical))
print(f"clean minus physical: {mean(a) - mean(b):+.3f} points a game, luck margin +/- {2 * sqrt(pstdev(a) ** 2 / len(a) + pstdev(b) ** 2 / len(b)):.3f}")
outshot = [r for r, l in zip(rows, labels) if l in (clean, physical)]
for label, group in (("all team-seasons", rows), ("the two outshot groups", outshot)):
print(f"physical against points, {label}: correlation {correlation([rough(r['z']) for r in group], [r['ppg'] for r in group]):.2f}")
# 5. are the groups stable? fit each half of the seasons on its own
for label, keep in (("2000/01-2012/13", lambda s: s < "1314"), ("2013/14-2025/26", lambda s: s >= "1314")):
_, half, _ = best_kmeans([r["z"] for r in rows if keep(r["season"])], 4)
print(f"{label}: " + "; ".join(f"dominance {dominance(c):+.2f}, physical {rough(c):+.2f}"
for c in sorted(half, key=lambda c: -dominance(c))))
# 6. a family tree of clubs with five or more seasons: join the two closest groups, again and again
clubs = defaultdict(list)
for r in rows:
clubs[r["team"]].append(r["z"])
style = {t: [mean(z[j] for z in zs) for j in range(8)] for t, zs in clubs.items() if len(zs) >= 5}
apart = lambda g, h: mean(sqrt(dist2(style[p], style[q])) for p in g for q in h) # average distance between members
groups = [[t] for t in sorted(style)]
while len(groups) > 1:
g, h = min(((g, h) for i, g in enumerate(groups) for h in groups[i + 1:]), key=lambda gh: apart(*gh))
print(f" {apart(g, h):.2f}: {' + '.join(g)} | {' + '.join(h)}")
groups = [x for x in groups if x not in (g, h)] + [g + h]
Further reading
- Visualizing K-Means Clustering, Naftali Harris: an interactive page where you place the starting centres and step through the assign-and-move loop yourself, including cases where it goes wrong.
- Clustering, scikit-learn user guide: the standard Python library's documentation, comparing K-means, hierarchical clustering and others on differently shaped data.
- An Introduction to Statistical Learning, by James, Witten, Hastie and Tibshirani: a free textbook whose unsupervised learning chapter covers K-means and hierarchical clustering, including how to choose the number of groups.