World Cup Final / Settled regulation-time audit
SETTLED / C4 MISSSpain 0-0 Argentina
The regulation-time result fell outside Turtle's locked score cloud. The four original predictions remain frozen below; the experiment is closed pending rebuild.
None landed. The actual 0-0 was modeled at 7.17%, ranked sixth, and left outside the four-node card.
Final laboratory record
Regulation time 0-0 C4 result: MISSTurtle failed its final exact-score test.
The model explicitly ranked 0-0 sixth at 7.17%, only 0.67 percentage points below its fourth published node. Its contradiction audit said the scoreless-draw branch survived the evidence floor, yet the selector gave that branch no place in C4.
This is a selection and calibration failure, not a missing-data excuse. Internal reproducibility did not establish real predictive reliability. The current protocol is closed pending a genuinely rebuilt and independently backtested method.
What was checked
The failed card was based on admitted data, not market prices.
The score model was completed before the evidence cutoff and reproduced twice. Those checks establish process consistency, not predictive accuracy.
Fifty regulation-time matches per team, source-cited and opponent-rated.
Clean sheets, BTTS, draws, margins, upsets, and high-score paths were tested.
Weight, rating, realm, lineup, and count-law variations produced no failed gate.
The independent A/B card files have the same bytes and the same four scores.
Human-facing data
What each team brought into the final.
Rates below describe regulation-time results available before the final. They are inputs, not promises.
Spain's current tournament defense is the strongest clean-sheet signal in the file. It is why 1-0 cannot be omitted.
Argentina scored in every admitted tournament match, while its recent defense allowed more than Spain's. That keeps 1-1 and 1-2 alive.
Comparable opponents
The model did not treat every old match equally.
It weighted matches against opponents near today's strength level. Spain had the stronger comparison sample; Argentina's was sparse and therefore pulled toward its broader history.
- Weighted GF
- 2.35
- Weighted GA
- 1.32
- Scored
- 96.7%
- Won
- 77.7%
- BTTS
- 57.0%
- 4+ goals
- 35.5%
- Adjusted GF
- 1.93
- Adjusted GA
- 0.77
- Scored
- 92.3%
- Won
- 65.3%
- BTTS
- 55.9%
- 4+ goals
- 20.5%
Sparse comparison: Bayesian shrinkage was applied. This is the largest data limitation.
Modelled match shape
Near-even result branches, with one to three goals carrying most of the mass.
Important: these tables summarize the same goal-count distribution. They are not independent votes added on top of it.
Selection failure
Why the four nodes were selected, and where that choice failed.
The model selected four high-ranked cells while preserving Spain's clean-sheet branch. It did not preserve the separate scoreless-draw branch that ultimately landed.
It is the single largest exact-score cell. Both sides have strong scoring evidence and the result model is close to even.
Argentina's long-run defense is strong, and an Argentina one-goal win is almost level with Spain's equivalent branch.
Spain conceded only 0.14 goals per match in seven current-tournament rows, so the clean-sheet override protects this node.
This is the strongest remaining cell after the draw and both clean-sheet paths. It preserves Argentina scoring twice without requiring a cushion.
| Rank | Score | Model probability | Status | Distance from fourth node |
|---|---|---|---|---|
| 1 | 1-1 | 11.65% | Locked | - |
| 2 | 0-1 | 9.12% | Locked | - |
| 3 | 1-0 | 9.11% | Locked | - |
| 4 | 1-2 | 7.84% | Locked | - |
| 5 | 2-1 | 7.81% | First score out | 0.03 percentage points |
| 6 | 0-0 | 7.17% | Outside C4 | 0.67 points |
| 7 | 0-2 | 6.11% | Outside C4 | 1.73 points |
| 8 | 2-0 | 6.08% | Outside C4 | 1.76 points |
Actual exclusion: 0-0 was sixth at 7.17%, only 0.67 percentage points below 1-2 Argentina. The branch passed its declared evidence floor but received no C4 node.
Stress tests
The selected nodes were internally stable and still wrong.
Forty-two model forks changed weights, ratings, opponent bands, lineup assumptions, and count laws. They tested persistence within the model, not calibration against unseen results.
Technical validation boundary
Data admission and computational mechanics passed: 100 model-valid rows, byte-identical twin cards, 21 contradiction attacks, and 42 passing stability forks. Exact historical copies of three execution-time source files were not preserved, so the old run cannot receive the strongest possible provenance certification. This page records a new public pre-kickoff freeze of the already-frozen model output; it does not backdate that publication.
Inspect the work
Data and model files.
Readers can inspect the rows and calculations behind the public card.