Guff AI ta ho niProject Turtle

World Cup Final / Settled regulation-time audit

SETTLED / C4 MISS

Spain 0-0 Argentina

The regulation-time result fell outside Turtle's locked score cloud. The four original predictions remain frozen below; the experiment is closed pending rebuild.

Final score
0-0 regulation
Public lock
2:24 PM ET
Actual rank
6th / 7.17%
Verdict
Experiment failed
Original regulation-time C4Unchanged after settlement
01 / Model leader1-111.65% model probability
02 / Argentina by one0-19.12% model probability
03 / Spain by one1-09.11% model probability
04 / Argentina by one1-27.84% model probability

None landed. The actual 0-0 was modeled at 7.17%, ranked sixth, and left outside the four-node card.

Final laboratory record

Regulation time 0-0 C4 result: MISS

Turtle failed its final exact-score test.

The model explicitly ranked 0-0 sixth at 7.17%, only 0.67 percentage points below its fourth published node. Its contradiction audit said the scoreless-draw branch survived the evidence floor, yet the selector gave that branch no place in C4.

This is a selection and calibration failure, not a missing-data excuse. Internal reproducibility did not establish real predictive reliability. The current protocol is closed pending a genuinely rebuilt and independently backtested method.

Failure class Scoreless-draw branch omittedPublic record Lock preserved unchangedProtocol state Closed pending rebuild
Read the complete audit

What was checked

The failed card was based on admitted data, not market prices.

The score model was completed before the evidence cutoff and reproduced twice. Those checks establish process consistency, not predictive accuracy.

100 / 100historical rows admitted

Fifty regulation-time matches per team, source-cited and opponent-rated.

21contradiction attacks

Clean sheets, BTTS, draws, margins, upsets, and high-score paths were tested.

42 / 42stress forks passed

Weight, rating, realm, lineup, and count-law variations produced no failed gate.

2 copiesidentical model output

The independent A/B card files have the same bytes and the same four scores.

Human-facing data

What each team brought into the final.

Rates below describe regulation-time results available before the final. They are inputs, not promises.

SpainStrength 10.00 / 10
Last 502.40 GF0.72 GA52% clean sheets
Last 101.70 GF0.30 GA70% clean sheets
World Cup, 7 matches1.86 GF0.14 GA86% clean sheets

Spain's current tournament defense is the strongest clean-sheet signal in the file. It is why 1-0 cannot be omitted.

ArgentinaStrength 9.83 / 10
Last 502.16 GF0.50 GA62% clean sheets
Last 102.50 GF0.60 GA50% clean sheets
World Cup, 7 matches2.14 GF0.86 GA100% scored

Argentina scored in every admitted tournament match, while its recent defense allowed more than Spain's. That keeps 1-1 and 1-2 alive.

Comparable opponents

The model did not treat every old match equally.

It weighted matches against opponents near today's strength level. Spain had the stronger comparison sample; Argentina's was sparse and therefore pulled toward its broader history.

Spain vs Argentina-level opponents10 matches / effective sample 6.45
Weighted GF
2.35
Weighted GA
1.32
Scored
96.7%
Won
77.7%
BTTS
57.0%
4+ goals
35.5%
Argentina vs Spain-level opponents5 matches / effective sample 3.25
Adjusted GF
1.93
Adjusted GA
0.77
Scored
92.3%
Won
65.3%
BTTS
55.9%
4+ goals
20.5%

Sparse comparison: Bayesian shrinkage was applied. This is the largest data limitation.

Modelled match shape

Near-even result branches, with one to three goals carrying most of the mass.

Result
Spain37.2%
Draw25.4%
Argentina37.4%
Total goals
0-125.4%
223.8%
321.3%
4+29.4%
Scoring state
BTTS53.8%
No BTTS46.1%
Spain clean26.6%
Argentina clean26.7%

Important: these tables summarize the same goal-count distribution. They are not independent votes added on top of it.

Selection failure

Why the four nodes were selected, and where that choice failed.

The model selected four high-ranked cells while preserving Spain's clean-sheet branch. It did not preserve the separate scoreless-draw branch that ultimately landed.

1-1
Model leader

It is the single largest exact-score cell. Both sides have strong scoring evidence and the result model is close to even.

0-1
Argentina clean-sheet route

Argentina's long-run defense is strong, and an Argentina one-goal win is almost level with Spain's equivalent branch.

1-0
Spain clean-sheet route

Spain conceded only 0.14 goals per match in seven current-tournament rows, so the clean-sheet override protects this node.

1-2
Three-goal Argentina route

This is the strongest remaining cell after the draw and both clean-sheet paths. It preserves Argentina scoring twice without requiring a cushion.

RankScoreModel probabilityStatusDistance from fourth node
11-111.65%Locked-
20-19.12%Locked-
31-09.11%Locked-
41-27.84%Locked-
52-17.81%First score out0.03 percentage points
60-07.17%Outside C40.67 points
70-26.11%Outside C41.73 points
82-06.08%Outside C41.76 points

Actual exclusion: 0-0 was sixth at 7.17%, only 0.67 percentage points below 1-2 Argentina. The branch passed its declared evidence floor but received no C4 node.

Stress tests

The selected nodes were internally stable and still wrong.

Forty-two model forks changed weights, ratings, opponent bands, lineup assumptions, and count laws. They tested persistence within the model, not calibration against unseen results.

1-1100%
0-195.8%
1-0100%
1-276.3%
Technical validation boundary

Data admission and computational mechanics passed: 100 model-valid rows, byte-identical twin cards, 21 contradiction attacks, and 42 passing stability forks. Exact historical copies of three execution-time source files were not preserved, so the old run cannot receive the strongest possible provenance certification. This page records a new public pre-kickoff freeze of the already-frozen model output; it does not backdate that publication.

Inspect the work

Data and model files.

Readers can inspect the rows and calculations behind the public card.