Monday is an evidence problem

Sunday offers a dangerous bargain: remember the good calls, explain away the bad ones, and arrive on Monday looking suspiciously undefeated. The Goblin declines. A forecast board earns its usefulness by preserving what it actually said, then attaching results to those exact claims.

This is a guide to our grading process, not a finished Week 1 performance review. Denver and Kansas City have yet to play as this edition is prepared on September 14. We will not turn an incomplete week into a victory lap or a funeral.

Grade the event, not the memory

Our current player forecasts define specific events: at least a stated number of rushing or receiving yards, or a scored rushing or receiving touchdown. A 60-yard threshold succeeds at exactly 60. A quarterback's passing touchdown does not satisfy a scoring-touchdown forecast. Those definitions must survive the final whistle.

A projected yardage total answers a different question from an event threshold. If the threshold is 60, the projection is 75 and the player finishes with 63, the event succeeds while the yardage estimate misses by 12. Keeping both measurements prevents a green checkmark from hiding the size of a forecasting error.

A player who participates and records zero does not disappear from the denominator. A genuine participation void needs evidence and an explanation. Blank extraction is not a zero, and an empty box-score table is not permission to improvise.

A real miss stays on the wall

The Melbourne report remains an uncomfortable exhibit. Atlas favored the Rams 27-23 at an uncalibrated 60%. San Francisco won 27-7. The published v2 grade preserves the losing winner forecast and records two successful player events out of five. It does not swap in a more flattering revision after seeing the score.

That single winner forecast has a Brier error of 0.36: square the difference between the published 0.60 probability and zero for the event that did not happen. A hypothetical 0.90 claim would have incurred 0.81. This illustrates why confidence should cost something when it is wrong. It does not prove that any percentage on this board is calibrated.

The older editorial v1 calls remain a separate history with null probabilities and their original quoted lines. Their results are not interchangeable with v2's independently chosen at-least thresholds. Combining the two would change the question after the answer arrived.

Frozen is not the same as fully checked

There are two different records to inspect. The kickoff cutoff blocks forecast revisions when the game starts. An explicit action lock also preserves a final bundle of injuries, inactive lists, weather and operator evidence. A forecast can be frozen by the first control even when the second bundle was not recorded.

Four Sunday afternoon games missed that explicit final-snapshot window. Their predictions remain frozen by the kickoff cutoff, but the missing evidence receipt remains missing. We will not backdate a completed scan. A checksum identifies a stored snapshot; it is not independent notarization against an administrator who changes the application.

Corrections append. Explanations wait for evidence.

Official final scores and player statistics are the basis for settlement. If an official correction changes a stat later, a new result must point to the previous result and explain the change. The forecast itself stays fixed. This keeps a correction distinguishable from a quiet rewrite.

After every game in the week is finished and every eligible report is graded, the review can report winner and player error, hit rates, yardage error and sample sizes. Until then, any running totals are partial. One explosive game is not enough to establish a new role; one failed touchdown call is not proof that the underlying opportunity vanished.

Useful learning begins with a testable question. Did we assume too many carries? Did a lineup change invalidate our protection argument? A proposed adjustment should be recorded as an experiment and judged on later, held-out evidence. A better explanation written on Monday is not proof of a better forecasting method, much less a newly trained language model. The receipts stay. Our ego can find another place to sit.