# Clue by clue. Naive Bayes, and the danger of counting twice

Source: https://www.footballdatascience.co.uk/learn/naive-bayes
Published: 2026-10-03

> Start with how often each result happens, then let every clue shift the odds. It's how spam filters work, and on three overlapping football clues it gets worse the more it's told. Merge the clues into one and it draws level with logistic regression.

**On the terraces:** A pundit who hears the same stat three ways round and counts it three times will sound very sure of himself. This piece shows a model making exactly that mistake, and the simple fix.

## The football question

Before kick-off you know three things about a match: which side is in better form, which has done better this season, and which did better last season. Each one is a clue. **What if you started with how often home sides win, and let each clue nudge the odds in turn?**

That's **naive Bayes**: [Bayes' theorem](/learn/bayes-theorem), applied one clue at a time. It's the method behind the first good spam filters, and it's quick, simple and often surprisingly good. On football it has a weakness worth knowing about.

## The concept

Start with the **base rates**: across the training matches, home wins 43.8%, draws 23.4%, away wins 32.8%. Then, for each result, ask how likely the clues are if that result is coming:

$$\begin{aligned} &P(\text{home} \mid \text{clues}) \\ &\quad \propto P(\text{home}) \times P(c_1 \mid \text{home}) \\ &\qquad \times P(c_2 \mid \text{home}) \times P(c_3 \mid \text{home}) \end{aligned}$$

<div class="plain" markdown="1">
In plain football

- **P(home)** is the base rate, how often home sides win before you know anything else.
- **P(c₁ | home)** asks: among matches the home side went on to win, how common was a form gap like this one? The same for this season's gap, **c₂**, and last season's, **c₃**.
- **∝** means "in proportion to". Work out the same product for a draw and an away win, then scale the three so they add up to 100%.
</div>

How common is a gap like this one? For each result, naive Bayes fits a bell curve, the [Normal distribution](/learn/normal-distribution), to the gaps in matches with that result. Home wins came with a this-season gap of +0.31 points a game on average, give or take 0.74; away wins with −0.39, give or take 0.74. A gap of +0.8 is far more typical of a home win than an away win, so it pushes the odds towards the home side.

The **naive** part is multiplying the clues together as if each were separate news. That's only fair if, once you know the result, the clues tell you nothing about each other. Spam filters get away with it: thousands of words, each a weak clue, none repeating another much. Football's three gaps are different.

The same three gaps as [logistic regression](/learn/logistic-regression), the same **3,900 training matches**, 2001/02 to 2020/21, and the same **990 test matches**, 2021/22 to 2025/26.

## Clue by clue

**Rangers 2–3 Motherwell, 26 April 2026.** Rangers had taken 2.40 points a game more than Motherwell over their last five, 0.45 more over the season and 0.68 more the season before. Here's naive Bayes adding one clue at a time. It's the test match where the two versions in this article disagree most, picked to show the problem clearly:

<figure class="rank-chart">
<div role="img" aria-label="The chance of a Rangers win, clue by clue. Base rates 44%. After recent form, 77%. After this season, 83%. After last season, 91%. With the three clues merged into one strength clue instead, 61%. Motherwell won 3–2.">

</div>
<figcaption>Each clue in turn pushes the naive forecast up, to 91%. One combined strength clue says 61%.</figcaption>
</figure>

Each step makes sense on its own. Together they don't: the hot form, the better season and the better last season are largely **the same news**, that Rangers are the stronger side, heard three times. Naive Bayes believes it three times and ends at 91%. Motherwell won.

## More clues, worse forecasts

Choose the clues honestly, on the held-back seasons 2016/17 to 2020/21, with the model fitted on 2001/02 to 2015/16:

| Clues | Held-back log loss |
|---|---|
| Base rates, no clues | 1.073 |
| Recent form only | 1.019 |
| Last season only | 0.987 |
| This season only | 0.983 |
| **All three** | **1.055** |

**Any one clue beats all three together.** All three barely beat knowing nothing. The reason is in how alike the clues are: on the training matches, the form and this-season gaps have a correlation of 0.76, this season and last season 0.73, form and last season 0.54. Three clues that overlap that much aren't three pieces of evidence.

The double-counting shows up as **overconfidence**. In the 160 test matches where all three clues made a home win 80% likely or more, they averaged **92%**. Home wins happened **80%** of the time.

## The fix: one clue, not three

If the clues overlap, merge them. Make one **strength clue**, a mix of the this-season and last-season gaps, and let naive Bayes use only that. How much of each? Again on the held-back seasons:

| This season / last season | Held-back log loss |
|---|---|
| 30% / 70% | 0.9741 |
| 40% / 60% | 0.9724 |
| **50% / 50%** | **0.9721** |
| 60% / 40% | 0.9729 |
| 70% / 30% | 0.9746 |

Half and half is best, and better than any single clue. It's the same idea as [PCA](/learn/pca-dimensionality-reduction), which boils overlapping stats down to the few directions that carry most of the information. Recent form is left out: it was the weakest clue alone, and it mostly repeats this season.

## Against the rest

The one look at the 990 test matches:

| Model | Test accuracy | Test log loss |
|---|---|---|
| Base rates | 47.2% | 1.056 |
| **Naive Bayes, three clues** | **52.8%** | **1.020** |
| **Naive Bayes, one strength clue** | **54.6%** | **0.951** |
| Logistic regression | 54.8% | 0.950 |
| KNN, 300 neighbours | 55.4% | 0.949 |
| Dixon-Coles | 55.3% | 0.945 |
| Bookmaker | 56.2% | 0.932 |

**With one clue, naive Bayes draws level with logistic regression**: 0.0006 behind, with a luck margin of ±0.0041. With three it falls most of the way back to knowing nothing. In the 160 sure-looking matches, the one-clue version averaged 76% against the 80% that happened, slightly cautious rather than badly overconfident.

The three-clue version even picks a draw as the likeliest result 10 times; logistic regression never does. The draw's bell curves are a little narrower than the others (spreads of 0.92, 0.68 and 0.56 points a game against 0.96, 0.74 and 0.61 for home wins), and multiplying three of them exaggerates that for matches near the middle.

Logistic regression doesn't have the overlap problem, because it fits all three weights together: when two features carry the same news, it splits the weight between them rather than counting it twice. That's why its recent-form weight came out close to zero. [K-nearest neighbours](/learn/nearest-neighbours) has the opposite weakness: it can't down-weight a feature at all.

## Why it matters

- **Check whether your features say the same thing.** Many models assume, somewhere, that features are separate evidence. Correlations between them are quick to check and often big.
- **Overconfidence is the symptom.** A model that's too sure, as in [calibration in depth](/learn/calibration-in-depth), is often counting something twice.
- **Simple models still earn a place.** With one well-built clue, a method from spam filtering matches the best simple model in the series, and you can see exactly how every number moves the odds.
- **It shines where clues are many and weak.** Text, with thousands of words each nudging the odds a little, is naive Bayes's home ground.

## Limitations

- **Bell curves are a choice.** The gaps within each result are roughly bell-shaped, but not exactly; the tails, where the mismatches are, fit least well.
- **One combined clue loses detail.** Merging the season gaps fixes the double count, but throws away any difference between this season and last that might matter.
- **The fix was chosen on one held-back block.** [Cross-validation](/learn/cross-validation) over several blocks would be steadier.
- **Three features, one league.** No injuries, team news or managers, and only the Scottish Premiership.

## Try it yourself

The [Naive Bayes Match Predictor](/models/naive-bayes) shows the odds moving clue by clue. Set the three gaps, then switch between three clues and one strength clue, and watch the overconfidence appear and disappear.

Or by hand: write down your home-win chance for this weekend's biggest mismatch. Now list every reason you're confident and ask how many of them are really the same reason. Would you knock a few points off?

## Reproduce the analysis

The results files are published by [football-data.co.uk](https://www.football-data.co.uk/scotlandm.php). Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as `SC0_2425.csv`; they aren't rehosted on this site. It runs in a few seconds:

```python
import csv
from collections import Counter, defaultdict
from datetime import datetime
from math import exp, log, pi, sqrt
from statistics import correlation

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
RESULTS = "HDA"
CLUES = ["recent form", "this season", "last season"]
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

def season(s):
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        games = [r for r in csv.DictReader(f) if r.get("FTR") in POINTS]
    games.sort(key=lambda r: datetime.strptime(r["Date"], "%d/%m/%Y" if len(r["Date"]) == 10 else "%d/%m/%y"))
    return games

def points_per_game(games):
    pts, n = Counter(), Counter()
    for r in games:
        for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
            pts[team] += p
            n[team] += 1
    return {t: pts[t] / n[t] for t in n}

# the same three gaps as part 8, home side minus away side, in points a game, all known before kick-off
rows = []
for s_last, s in zip(names, names[1:]):
    last = points_per_game(season(s_last))
    promoted = 0.85 * sum(last.values()) / len(last)
    history = defaultdict(list)
    for r in season(s):
        h, a = r["HomeTeam"], r["AwayTeam"]
        if len(history[h]) >= 5 and len(history[a]) >= 5:
            rows.append({"season": s, "match": f"{h} {r['FTHG']}-{r['FTAG']} {a}", "date": r["Date"], "y": r["FTR"], "x": [
                (sum(history[h][-5:]) - sum(history[a][-5:])) / 5,
                sum(history[h]) / len(history[h]) - sum(history[a]) / len(history[a]),
                last.get(h, promoted) - last.get(a, promoted)]})
        for team, p in zip((h, a), POINTS[r["FTR"]]):
            history[team].append(p)

def with_strength(data, w):  # one combined clue: w of this season's gap plus (1 - w) of last season's
    for r in data:
        r["x"] = r["x"][:3] + [w * r["x"][1] + (1 - w) * r["x"][2]]

def naive_bayes(fit, clues):
    """Base rates, then for each clue a bell curve per result: its average and spread among home wins, draws, away wins."""
    prior = Counter(r["y"] for r in fit)
    curves = {}
    for c in RESULTS:
        xs = [r["x"] for r in fit if r["y"] == c]
        for j in clues:
            m = sum(x[j] for x in xs) / len(xs)
            curves[c, j] = (m, sqrt(sum((x[j] - m) ** 2 for x in xs) / len(xs)))
    def forecast(x, upto=None):  # upto: only the first few clues, to watch the odds move clue by clue
        score = {c: log(prior[c] / len(fit)) for c in RESULTS}
        for j in clues[:upto]:
            for c in RESULTS:
                m, sd = curves[c, j]
                score[c] += -log(sd * sqrt(2 * pi)) - (x[j] - m) ** 2 / (2 * sd ** 2)  # log of the bell curve's height
        top = max(score.values())
        e = {c: exp(v - top) for c, v in score.items()}
        return {c: e[c] / sum(e.values()) for c in RESULTS}
    return forecast, prior, curves

def logistic(fit):  # part 8's model, standardised and refitted here, for a match-by-match comparison
    mean = [sum(r["x"][j] for r in fit) / len(fit) for j in range(3)]
    sd = [sqrt(sum((r["x"][j] - mean[j]) ** 2 for r in fit) / len(fit)) for j in range(3)]
    z = lambda x: [(x[j] - mean[j]) / sd[j] for j in range(3)]
    def predict(w, v):
        e = {c: exp(w[c][0] + sum(a * b for a, b in zip(w[c][1:], v))) for c in RESULTS}
        return {c: e[c] / sum(e.values()) for c in RESULTS}
    w = {c: [0.0] * 4 for c in RESULTS}
    data = [(z(r["x"]), r["y"]) for r in fit]
    for _ in range(200):
        grad = {c: [0.0] * 4 for c in "HA"}
        for v, y in data:
            p = predict(w, v)
            for c in "HA":
                grad[c] = [g + (p[c] - (c == y)) * t for g, t in zip(grad[c], [1] + v)]
        for c in "HA":
            w[c] = [a - g / len(data) for a, g in zip(w[c], grad[c])]
    return lambda x: predict(w, z(x))

def log_loss(model, data):
    return sum(-log(model(r["x"])[r["y"]]) for r in data) / len(data)

def base_rates(fit):
    n = Counter(r["y"] for r in fit)
    return lambda x: {c: n[c] / len(fit) for c in RESULTS}

fit = [r for r in rows if r["season"] < "1617"]
held = [r for r in rows if "1617" <= r["season"] < "2122"]
train = [r for r in rows if r["season"] < "2122"]
test = [r for r in rows if r["season"] >= "2122"]
pairs = ((0, 1), (0, 2), (1, 2))
print("how alike the clues are (training matches): " + ", ".join(
    f"{CLUES[i]} and {CLUES[j]} {correlation([r['x'][i] for r in train], [r['x'][j] for r in train]):.2f}" for i, j in pairs))

# 1. on the held-back seasons: which clues, and how to combine the two season gaps
print(f"held back: base rates {log_loss(base_rates(fit), held):.4f}")
for clues in ([0, 1, 2], [0], [1], [2]):
    print(f"  naive Bayes on {', '.join(CLUES[j] for j in clues)}: {log_loss(naive_bayes(fit, clues)[0], held):.4f}")
scores = {}
for w in (0.3, 0.4, 0.5, 0.6, 0.7):
    with_strength(fit + held, w)
    scores[w] = log_loss(naive_bayes(fit, [3])[0], held)
    print(f"  one strength clue, {w:.0%} this season: {scores[w]:.4f}")
best = min(scores, key=scores.get)
print(f"chosen: {best:.0%} this season, {1 - best:.0%} last season")

# 2. refit on all training matches; score the test seasons once
with_strength(rows, best)
naive, prior, curves = naive_bayes(train, [0, 1, 2])
fixed, _, strength = naive_bayes(train, [3])
lr = logistic(train)
for name, model in (("base rates", base_rates(train)), ("naive Bayes, three clues", naive), ("naive Bayes, one strength clue", fixed), ("logistic regression", lr)):
    right = sum(max(RESULTS, key=model(r["x"]).get) == r["y"] for r in test)
    picks = Counter(max(RESULTS, key=model(r["x"]).get) for r in test)
    print(f"test: {name}: log loss {log_loss(model, test):.4f}, accuracy {right / len(test):.1%}, picks {dict(picks)}")
gap = [log(lr(r["x"])[r["y"]]) - log(fixed(r["x"])[r["y"]]) for r in test]
avg = sum(gap) / len(gap)
print(f"one strength clue minus logistic regression: {avg:+.4f}, luck margin +/- {2 * sqrt(sum((g - avg) ** 2 for g in gap) / (len(gap) - 1) / len(gap)):.4f}")

# 3. counting the same news twice: test matches where three clues make a home win 80% or more
sure = [r for r in test if naive(r["x"])["H"] >= 0.8]
print(f"three clues say home 80%+ in {len(sure)} test matches: they average {sum(naive(r['x'])['H'] for r in sure) / len(sure):.0%}, "
      f"one strength clue {sum(fixed(r['x'])['H'] for r in sure) / len(sure):.0%}, home wins happened {sum(r['y'] == 'H' for r in sure) / len(sure):.0%}")
even = [0, 0, 0, 0]
print("even game: three clues", {c: f"{p:.1%}" for c, p in naive(even).items()}, "| one clue", {c: f"{p:.1%}" for c, p in fixed(even).items()})

# 4. clue by clue, for the test match where the two disagree most about a home win
r = max(test, key=lambda r: abs(naive(r["x"])["H"] - fixed(r["x"])["H"]))
print(f"{r['match']} ({r['date']}), gaps " + ", ".join(f"{v:+.2f}" for v in r["x"][:3]) + f", strength {r['x'][3]:+.2f}")
for k in range(4):
    p = naive(r["x"], upto=k)
    print(f"  naive, after {k} clue{'s' if k != 1 else ''}: home {p['H']:.0%}, draw {p['D']:.0%}, away {p['A']:.0%}")
p = fixed(r["x"])
print(f"  one strength clue: home {p['H']:.0%}, draw {p['D']:.0%}, away {p['A']:.0%}")

# 5. what the model learned, for the Naive Bayes Match Predictor
print("base rates:", {c: round(prior[c] / len(train), 4) for c in RESULTS})
for j, label in enumerate(CLUES + ["strength"]):
    source = strength if j == 3 else curves
    print(f"{label}:", {c: tuple(round(v, 4) for v in source[c, j]) for c in RESULTS})
```

## Further reading

- [A Plan for Spam](https://www.paulgraham.com/spam.html), Paul Graham, August 2002: the essay that popularised Bayesian spam filtering, with the word-by-word probabilities it combines.
- [Naive Bayes, scikit-learn user guide](https://scikit-learn.org/stable/modules/naive_bayes.html): the standard Python library's documentation, including the Gaussian version used here.
- [An Introduction to Statistical Learning](https://www.statlearning.com/), by James, Witten, Hastie and Tibshirani: a free textbook whose classification chapter covers naive Bayes alongside logistic regression and KNN.
