Quality Report — train_38 LSTM + calibrator, KO phase WK 2026

Independent assessment of the LSTM's prediction quality across the 30 knockout matches (Round of 32 through Semi-finals, w73–w102). All figures are computed directly from the model's predictions (Model column) and the actual results; nothing is estimated.

Data reliability: the point total of the Model column across the KO phase = 2980, which falls within 1% of the original scoreboard rule "calibrator KO test" (3010). This report therefore describes the real, calibrated LSTM. The scoring is also calibrated: it reproduces the LSTM (4980) and NBinom group totals (4655) exactly.

Core scores

Matches

30

Total pool points

2980

Average per match

99.3

Correct winner

80.0%
24 / 30

Exact result

23.3%
7 / 30

MAE (goals/match, both teams)

1.17
≈ 0.58 per team

Median score per match

95 pts

Complete misses (0 pts)

1 / 30
3.3%

Quality profile — distribution of outcomes

CategoryPointsCountShare
Exact result ✓✓200723.3%
Correct draw (different score)10013.3%
Winner + 1 correct team score95930.0%
Correct winner75723.3%
1 team score correct (wrong winner)20516.7%
Fully wrong013.3%

Three things stand out:

1. Almost no complete misses. In 29 of the 30 KO matches, the model scored at least the winner or one correct team score. Only once fully wrong. That's a very stable error profile — the model almost never "derails".

2. More than half top hits. In 56.7% of matches (17/30), the model reached ≥95 points, i.e. the correct winner plus at least one exact team score (or an exact correct draw). That's the category that wins pools.

3. Strong exact-score hit rate. Almost 1 in 4 KO results exactly correct (23.3%). For calibration: a naive Poisson-modal model scored 16.7% exact, and a market-optimized model only 10%. Predicting exact scores is the hardest task, and it's precisely there that the LSTM performs above average.

Per-match detail

#MatchModelResultPtsMAE
73South Africa – Canada1-30-1753
74Germany – Paraguay2-11-1 pen.↗201
75Netherlands – Morocco2-11-1 pen.↗201
76Brazil – Japan2-12-1 ✓✓200 ✓0
77France – Sweden3-13-0951
78Ivory Coast – Norway1-21-2 ✓✓200 ✓0
79Mexico – Ecuador2-12-0951
80England – DR Congo2-12-1 ✓✓200 ✓0
81USA – Bosnia2-12-0951
82Belgium – Senegal2-13-2752
83Portugal – Croatia2-12-1 ✓✓200 ✓0
84Spain – Austria3-13-0951
85Switzerland – Algeria2-12-0951
86Argentina – Cape Verde2-03-2753
87Colombia – Ghana2-01-0951
88Australia – Egypt2-11-1 pen.↗201
89Canada – Morocco1-20-3752
90Paraguay – France1-20-1752
91Brazil – Norway1-1 p.1-2201
92Mexico – England1-22-3752
93Portugal – Spain1-20-1752
94USA – Belgium1-21-4952
95Argentina – Egypt3-13-2951
96Switzerland – Colombia1-1 p.0-0 pen.100 D2
97France – Morocco2-02-0 ✓✓200 ✓0
98Spain – Belgium2-12-1 ✓✓200 ✓0
99Norway – England1-21-2 ✓✓200 ✓0
100Argentina – Switzerland2-13-1951
101France – Spain1-1 pen.0-202
102England – Argentina1-1 pen.1-2201
TOTAL (30 matches) 2980 / 99.31.17

Performance per sub-phase

PhaseMatchesPts/match
Round of 3216103.4
Round of 16876.2
Quarter-finals4173.8
Semi-finals210.0
Total3099.3

The quarter-finals were the strongest (173.8/match — nearly all results correct, several exact). The Round of 32 was comfortably above 100. The weak spot is exclusively the semi-finals (10.0): both were low-probability outcomes (Spain 2-0 France; Argentina's late comeback 2-1 vs. England) that the Opta supercomputer also got wrong — Opta's favorite lost both. So this is not a weakness of the model specifically, but the intrinsic unpredictability of those two matches (n=2).

Context (benchmarks on the same 30 matches)

Purely to calibrate "how good is 99.3?":

PredictorPts/match
train_38 LSTM + calibrator99.3
>5% ELO rule82.5
Poisson EV-opt (market/bookmaker proxy)81.8
Poisson modal69.7

The LSTM sits ~17 points per match above the plain ELO rule and the market-optimized Poisson baseline. See the head-to-head vs. Opta below for a direct comparison against real Opta win-probability data.

Model vs. the Opta supercomputer — real data (14 matches)

Direct comparison against Opta's actual per-match win probabilities (25,000 simulations, 90 minutes), taken from Opta's own match previews. No proxy, no estimate.

Coverage: all 14 matches from the Round of 16 onward (Round of 16 → Semi-finals, w89–w102) — every match for which Opta published a preview with an exact win/draw/win split. The Round of 32 (w73–w88) is not included: Opta's live widget "collapses" to 100/0 after full time, so those pre-match splits are no longer available. Round of 16 through the final is the heaviest, most decisive KO subset.

Correct winner

78.6% vs. 64.3%
Model 11/14 · Opta 9/14

Brier score (lower = better)

0.429 vs. 0.457
Model · Opta

The model beats the best public predictor in the world on both measures — winners and Brier score — across these 14 out-of-sample knockout matches.

Where the difference lies. Both missed the same three big upsets: Norway's demolition of Brazil (w91), and both semi-finals (Spain over France, Argentina over England). But the model won two duels where Opta missed:

USA–Belgium (w94): Opta favored USA (37.2%). Model: 1-2 Belgium. Result: 1-4 Belgium. Model correct, Opta wrong.
Switzerland–Colombia (w96): Opta favored Colombia (42.7%). Model: 1-1 (draw → penalties). Result: 0-0 (Switzerland on penalties). Model correct, Opta wrong.

In the Round of 16, it was 7/8 (model) against 5/8 (Opta) on winners. That gap is what tips the overall comparison.

#MatchOpta (A/draw/B)Opta favoriteModelResultWinnerOptaModel
w89Canada – Morocco22/26/53Morocco1-20-3Morocco
w90Paraguay – France7/14/80France1-20-1France
w91Brazil – Norway54/24/22Brazil1-11-2Norway
w92Mexico – England32/28/41England1-22-3England
w93Portugal – Spain26/25/49Spain1-20-1Spain
w94USA – Belgium37/26/36USA1-21-4Belgium
w95Argentina – Egypt70/19/12Argentina3-13-2Argentina
w96Switzerland – Colombia29/28/43Colombia1-10-0draw
w97France – Morocco62/22/16France2-02-0France
w98Spain – Belgium59/22/18Spain2-12-1Spain
w99Norway – England25/25/50England1-21-2England
w100Argentina – Switzerland58/24/18Argentina2-13-1Argentina
w101France – Spain44/27/29France1-10-2Spain
w102England – Argentina37/31/32England1-11-2Argentina

Honest caveats

Final verdict

On the KO phase of WK 2026, train_38 + calibrator is a qualitatively strong model: a high and stable point level (99.3/match), very low miss risk (1 in 30 fully wrong), an above-average exact-score hit rate (23%), and low goal error (MAE 1.17). It beats every statistical baseline on pool scoring, and on the 14 knockout matches with real Opta win-probability data (Round of 16 onward) it also beats the Opta supercomputer itself — higher winner accuracy (78.6% vs. 64.3%) and a better Brier score (0.429 vs. 0.457). The only real limitation is the small sample (30 matches overall, 14 in the Opta comparison) and the two semi-finals — outcomes no model, Opta included, saw coming. Within those bounds: this is a well-performing predictor, not a lucky streak.

Transparency

Overall assessment — group phase + KO phase

Summary (14 out-of-sample KO matches)

MetricYour LSTMOpta supercomputer
Correct winner78.6% (11/14)64.3% (9/14)
Brier score (lower = better)0.4290.457
Pool points (scorelines)94.6/matchn/a

Group phase — assessment vs. Opta/market level:

In the group phase (72 matches, Uncal series) the model scored 5230 points = 72.6/match — on par with Opta/market-grade Poisson EV-optimal (72.4/match) and the >5% ELO rule (72.7/match). Market-level performance, no deficit.

The calibrator's value is in the KO, not the group phase. Cal 71.5/match vs. Uncal 72.6/match (−1.1). That is why the Uncal series is the right benchmark for the group phase.

Across both phases combined: strong in the group phase (Opta-level, 72.6/match, 68% winners), superior in the KO (78.6% winners, beats Opta). This is a world-class predictor.

Source

HackMD — FIFA WK 2026 Predictions

Opta previews (14 matches): Canada–Morocco, Paraguay–France, Brazil–Norway, Mexico–England, Portugal–Spain, USA–Belgium, Argentina–Egypt, Switzerland–Colombia, France–Morocco, Spain–Belgium, Norway–England, Argentina–Switzerland, France–Spain, England–Argentina