Skip to content

What happens next season? Time series and ARIMA

A club's points a game, season after season, is a time series. For most clubs a season remembers the last two or three; a match remembers nothing. Smoothing over every past season forecasts next season better than last season alone, about nine points either way.

Intermediate Part 19 of Machine Learning Through Football

New to the notation? The symbols explained

Contents

The football question

Hearts took 80 points in 2025/26, 2.11 a game, their best season since they came back up in 2021/22. How many should they expect in 2026/27? The same again? Something closer to their usual? And does a team that's winning carry that into the next match?

Questions about how a number moves over time are time series questions, and the classic family of models for them is ARIMA. Most of the models in this series looked across matches at one moment. This one looks along time.

The concept

A time series is a run of numbers in order: here, a club's points a game season after season, or its points match after match. The first thing to ask of one is how much it remembers: how closely each value matches the one before, the one before that, and so on. That's its autocorrelation.

ARIMA puts three ideas together, one per letter:

  • AR, autoregressive: next season is built from the last few seasons, each with its own weight.
  • I, integrated: work with the changes from season to season rather than the levels, when the level itself wanders.
  • MA, moving average: next season is nudged by the last few surprises, how far each season beat or missed what was expected.

You'll have met the simplest one already. Rolling downhill fitted this line:

$$\begin{aligned} &\text{this season} = \text{average} \\ &\quad + w \times (\text{last season} - \text{average}) \end{aligned}$$

In plain football

  • Average is the typical side's points a game, 1.38 here.
  • w says how much of last season's gap from average carries over. Below 1 means part of every good or bad season is luck that won't repeat: regression to the mean.
  • This is an AR(1) model: autoregressive, one season back.

The data: points a game for every Scottish Premiership team-season from 2000/01 to 2025/26, 312 team-seasons. A forecast needs at least the season before, so it's made for each team-season whose club was also in the Premiership the season before. Choices are made on the held-back seasons 2016/17 to 2020/21, and the test is 2021/22 to 2025/26, 53 forecasts.

How much does a season remember?

Line up every club's points a game with its points one season earlier, two seasons earlier, and so on, and measure the correlation:

Across all clubs, a season still matches the one five years earlier at 0.71. Take out Celtic and Rangers and the memory fades within a few seasons.

Across all clubs it barely fades: five seasons back still correlates at 0.71. But that's mostly two clubs. Celtic and Rangers are far above everyone else every season, and their gap alone makes any season look like any other. Take them out and the rest remember much less: 0.46 one season back, 0.25 three back, 0.07 five back. For most clubs, how good they are drifts, and after a few seasons it has drifted a long way. A level that drifts rather than snapping back to one fixed value is what the I in ARIMA is for: model the changes, not the level.

It also means one season back isn't all the information. Fit an AR(2), with two seasons back, and last season gets a weight of 0.59 and the season before 0.35. The season before last still counts for over a third.

Smoothing: every past season, the latest counting most

Rather than choose how many seasons back to look, exponential smoothing uses them all, each counting a little less than the one after it. Start at the league average, then after every season move part of the way towards what the club actually did:

$$\begin{aligned} &\text{new level} = \text{old level} \\ &\quad + \alpha \times (\text{season} - \text{old level}) \end{aligned}$$

In plain football

  • Level is the model's view of how good the club really is. The forecast for next season is simply the latest level.
  • α, alpha, is how far each season moves it. Alpha 1 means "only last season counts"; small alpha means "change your mind slowly".
  • Season − old level is the surprise: how far the club beat or missed what was expected.

That's an ARIMA(0,1,1) model in another form: it works with changes (the I) and is driven by the latest surprise (the MA). It's also the shape of Elo's update: a rating moves part of the way towards what each result says. Chosen on the held-back seasons, the best alpha is 0.70, so each new season counts for most of the level but never all of it.

Here's Hearts since they came back up in 2021/22:

Dots are what Hearts did; the dashed line is the smoothed level. After 2.11 in 2025/26 the level rises to 1.91, not all the way, because the seasons before were lower. The shaded band is the forecast's typical miss.

Their 2.11 in 2025/26 lifts the level from 1.46 to 1.91, not to 2.11, because the four seasons before it were between 1.37 and 1.79. The forecast for 2026/27 is 1.91 points a game, about 73 points over 38 matches, against the 80 they took.

The honest test

Forecast each of the 53 test team-seasons, 2021/22 to 2025/26, using only seasons before it:

Forecast Typical miss, points a season
League average 18.9
Same as last season 10.1
AR(1), w 0.87 9.8
Smoothing, alpha 0.70 9.0

The typical miss is the square root of the average squared miss, in points a game, times 38.

Smoothing wins. Its edge over "same as last season" is clear of luck: a squared miss 0.0148 smaller, with a luck margin of ±0.0097. Over the AR(1) line, which on these training seasons comes out at w 0.87 (rolling downhill found 0.83 on every season), its edge is 0.0109, right at the edge of its luck margin of ±0.0108. A forecast built on several seasons beats one built on the last season alone, which is what naive Bayes and logistic regression found from the other direction: two seasons tell you more than one.

And nine points either way is the honest size of the uncertainty. Before a ball is kicked, a season forecast from results alone can't do much better.

Match to match: no momentum

Now the same question along a season: does a team's result in one match say anything about the next, once you know how good the team was that season? Take every pair of consecutive matches, 11,446 of them, and measure the correlation between the two results, each measured from that team's season average.

It's −0.025. That looks like a tiny bounce-back, but removing each season's average pulls the correlation down by about 0.027 on its own, even if every result were independent. So the honest reading is no momentum at all. A team on a winning run is a good team, not a team carried by the run, which is what the form myth found too.

Why it matters

  • Seasons remember, matches don't. For most clubs a season carries over for two or three years, far longer for the Old Firm; last Saturday's result tells you nothing about next Saturday beyond what the season already does.
  • Use more than the latest number. Smoothing over every past season beat last season alone, here and in most forecasting.
  • Many models are time series models in disguise. Elo's rating update, the decay that weights recent matches in Dixon-Coles, regression to the mean: all versions of "move part of the way towards what you just saw".
  • Say how wrong it'll be. "About 73 points, usually within nine" is a forecast; "73 points" is a guess.

Limitations

  • Short series. A club has at most 26 seasons here, and fewer unbroken. Full ARIMA, with its orders chosen from the data, needs long runs: monthly attendances, weekly goals across a league. With this little data the models have to stay small.
  • Breaks. A new owner, a new manager, administration: a club's level can jump, and smoothing takes a season or two to catch up.
  • Promoted sides have no run. A club just promoted has no Premiership season before, so these forecasts can't start; in their first season back, promoted sides have averaged 1.15 points a game, about 44 points.
  • Results only. Signings, sales and budgets, which drive next season, aren't in the data.

Try it yourself

The Next Season Forecaster does this for every club in the 2026/27 Premiership: its seasons as a chart, the smoothed level, and next season's forecast in points with its range.

Or by hand: take your club's points from the last three seasons. Weight them 0.70, 0.21 and 0.06 (newest first), add 0.03 times 52 points, the league average over 38 games, and you have a rough smoothed forecast. How far is it from last season's total?

Reproduce the analysis

The results files are published by football-data.co.uk. Download the Premiership file (SC0) for each season from 2000/01 to 2025/26 and save each under its own name, such as SC0_2425.csv; they aren't rehosted on this site. It runs in under a second:

Show the Python111 lines, ready to copy and run.
import csv
from collections import Counter, defaultdict
from math import sqrt
from statistics import correlation, mean

POINTS = {"H": (3, 0), "D": (1, 1), "A": (0, 3)}
names = [f"{y % 100:02d}{(y + 1) % 100:02d}" for y in range(2000, 2026)]

# points a game for every Premiership team-season, and each team's points match by match
ppg, games = {}, defaultdict(list)
for s in names:
    with open(f"SC0_{s}.csv", encoding="latin-1") as f:
        for r in csv.DictReader(f):
            if r.get("FTR") in POINTS:
                for team, p in zip((r["HomeTeam"], r["AwayTeam"]), POINTS[r["FTR"]]):
                    games[s, team].append(p)
for key, ps in games.items():
    ppg[key] = sum(ps) / len(ps)

def run_to(i, team):  # the team's unbroken run of Premiership seasons up to names[i], oldest first
    run = []
    while i >= 0 and (names[i], team) in ppg:
        run.insert(0, ppg[names[i], team])
        i -= 1
    return run

def history(s, team):  # the run before season s
    return run_to(names.index(s) - 1, team)

cases = [{"season": s, "team": t, "actual": v, "past": history(s, t)} for (s, t), v in ppg.items() if history(s, t)]
fit = [c for c in cases if c["season"] < "1617"]
held = [c for c in cases if "1617" <= c["season"] < "2122"]
train = [c for c in cases if c["season"] < "2122"]
test = [c for c in cases if c["season"] >= "2122"]
AVERAGE = mean(v for (s, _), v in ppg.items() if s < "2122")
print(f"{len(ppg)} team-seasons; {len(fit)} / {len(held)} / {len(test)} forecasts to fit / hold back / test; "
      f"average {AVERAGE:.2f} points a game")

# 1. how much does a season remember? correlation with 1 to 5 seasons before, with and without Celtic and Rangers
for lag in range(1, 6):
    line = []
    for keep in (lambda t: True, lambda t: t not in ("Celtic", "Rangers")):
        pairs = [(c["past"][-lag], c["actual"]) for c in cases if len(c["past"]) >= lag and keep(c["team"])]
        line.append(f"{correlation(*zip(*pairs)):.2f} ({len(pairs)} pairs)")
    print(f"  {lag} season(s) back: all clubs {line[0]}, without Celtic and Rangers {line[1]}")

# 2. forecasting next season
def ar1(data):  # this season = average + w x (last season - average): least squares for w
    x = [c["past"][-1] - AVERAGE for c in data]
    y = [c["actual"] - AVERAGE for c in data]
    w = sum(a * b for a, b in zip(x, y)) / sum(a * a for a in x)
    return lambda past: AVERAGE + w * (past[-1] - AVERAGE), w

def ar2(data):  # the same, with the season before last too: least squares for two weights
    rows = [(c["past"][-1] - AVERAGE, c["past"][-2] - AVERAGE, c["actual"] - AVERAGE) for c in data if len(c["past"]) >= 2]
    s11 = sum(a * a for a, _, _ in rows); s22 = sum(b * b for _, b, _ in rows); s12 = sum(a * b for a, b, _ in rows)
    s1y = sum(a * y for a, _, y in rows); s2y = sum(b * y for _, b, y in rows)
    det = s11 * s22 - s12 ** 2
    return (s1y * s22 - s2y * s12) / det, (s2y * s11 - s1y * s12) / det

def smooth(alpha):  # exponential smoothing, ARIMA(0,1,1): start at the average, move alpha of the way to each season
    def forecast(past):
        level = AVERAGE
        for v in past:
            level += alpha * (v - level)
        return level
    return forecast

def typical_miss(model, data):  # root mean squared miss, points a game
    return sqrt(mean([(model(c["past"]) - c["actual"]) ** 2 for c in data]))

by_alpha = {a / 20: typical_miss(smooth(a / 20), held) for a in range(1, 21)}
alpha = min(by_alpha, key=by_alpha.get)
print(f"held back: smoothing alpha {alpha:.2f} best ({by_alpha[alpha]:.3f}); alpha 1, just last season: {by_alpha[1.0]:.3f}")
first, second = ar2(train)
print(f"AR(2) weights: last season {first:.2f}, the season before {second:.2f}")
model_ar1, w = ar1(train)
models = {"league average": lambda past: AVERAGE, "last season": lambda past: past[-1],
          f"AR(1), w {w:.2f}": model_ar1, f"smoothing, alpha {alpha:.2f}": smooth(alpha)}
for name, model in models.items():
    miss = typical_miss(model, test)
    print(f"test: {name}: misses by {miss:.3f} points a game, {38 * miss:.1f} points a season")
for rival in ("last season", f"AR(1), w {w:.2f}"):
    gap = [(smooth(alpha)(c["past"]) - c["actual"]) ** 2 - (models[rival](c["past"]) - c["actual"]) ** 2 for c in test]
    g = mean(gap)
    print(f"  smoothing minus {rival}, squared miss: {g:+.4f}, luck margin +/- {2 * sqrt(sum((x - g) ** 2 for x in gap) / (len(gap) - 1) / len(gap)):.4f}")

# 3. momentum: does one match's result say anything about the next, once the season's level is known?
xs, ys, bias = [], [], []
for ps in games.values():
    m = mean(ps)
    xs += [p - m for p in ps[:-1]]
    ys += [p - m for p in ps[1:]]
    bias.append(-1 / (len(ps) - 1))  # what removing the season's average does on its own, with no momentum at all
print(f"match to match: correlation {correlation(xs, ys):+.3f} over {len(xs)} pairs; pure chance gives about {mean(bias):+.3f}")

# 4. promoted sides: points a game in their first season back
promoted = [v for (s, t), v in ppg.items() if s != names[0] and (names[names.index(s) - 1], t) not in ppg]
print(f"promoted sides' first season: {mean(promoted):.2f} points a game ({len(promoted)} team-seasons)")

# 5. forecasts for 2026/27, smoothing over each club's run of seasons up to 2025/26
for team in ("Celtic", "Hearts", "Aberdeen", "Dundee"):
    past = run_to(len(names) - 1, team)
    f = smooth(alpha)(past)
    print(f"2026/27 {team}: {f:.2f} points a game, about {38 * f:.0f} points; 2025/26 was {past[-1]:.2f}")

# 6. Hearts season by season since they came back up, with the smoothed level after each season
run, level = run_to(len(names) - 1, "Hearts"), AVERAGE
for s, v in zip(names[-len(run):], run):
    level += alpha * (v - level)
    print(f"  Hearts 20{s[:2]}/{s[2:]}: {v:.2f}, level {level:.2f}")

Further reading

  • Forecasting: Principles and Practice, by Rob Hyndman and George Athanasopoulos: a free online textbook, the standard introduction to time series forecasting, with chapters on exponential smoothing and ARIMA.
  • Introduction to ARIMA models, Robert Nau, Duke University: lecture notes that explain the AR, I and MA parts in plain terms, including why exponential smoothing is an ARIMA(0,1,1) model.
  • ARIMA, statsmodels documentation: the Python library most people use to fit full ARIMA models.

Get the weekly email

One email a week, on Fridays, with everything new, and the occasional note. Unsubscribe in one click. How your email is used.