Every model idea gets a pass/fail rule written down before the data that tests it exists. Failures stay on the record.
kalman adds information beyond the opening line rejected
Question. Encompassing actual = a + b*open + c*kalman has c > 0.
Result. Margin TEST c = +0.258 (p 0.019), VALIDATE c = -0.101 (p 0.27). Total TEST +0.047 (p 0.66), VALIDATE -0.026.
Verdict. FAIL: margin sign flips between periods, the same signature as the dead featured-totals candidate. Not an edge.
Registered 2026-10-02 · resolved 2026-10-02
Kalman vs decay_od_div (which recency method) kalman wins margin, decay_od_div wins totals
Question. kalman beats decay_od_div on paired residual sd.
Result. Margin +0.69 TEST / +0.38 VALIDATE; totals -0.36 / -0.34.
Verdict. PASS on margin; decay_od_div remains the better totals model.
Registered 2026-10-02 · resolved 2026-10-02
Kalman dynamic ratings beat elo beats elo on margin, not totals
Question. kalman (TUNE-selected sigma_obs/q_week/q_season/rho) beats elo on paired residual sd.
Result. Margin: TEST 15.44 vs 16.54 (+1.10), VALIDATE 15.92 vs 16.69 (+0.77). Totals: -0.28 / -0.25. TEST closing line margin 15.05.
Verdict. PASS on margin. Best margin model in the repo, 0.39 behind the TEST closing line. Worse than elo on totals.
Registered 2026-10-02 · resolved 2026-10-02
Forward 2026 wk6+: decay_od_div and 50/50 elo blend vs elo, and vs the close open
Question. F1 decay_od_div beats elo; F2 fixed 0.5/0.5 blend of elo and decay_od_div beats elo; F3 either one has encompassing c > 0 against the ESPN closing line.
Registered 2026-10-02
decay_od_div adds information beyond the line rejected
Question. Encompassing c > 0 against the closing line (HOLDOUT) and the opening line (HOLDOUT-2).
Result. HOLDOUT close: margin c = -0.010 (p 0.88), total c = -0.205 (p 0.12). HOLDOUT-2 open: margin -0.099, total +0.130 (p 0.35).
Verdict. FAIL. Encompassed.
Registered 2026-10-02 · resolved 2026-10-02
decay_od_div beats elo rejected
Question. The fixed model beats elo on paired residual sd (H1 re-asked on new seasons).
Result. HOLDOUT margin +0.11, total 0.00; HOLDOUT-2 margin -0.10, total +0.08.
Verdict. FAIL. Level with elo historically. Best rating model in 2026 wk1-4 (exploratory): 16.66 vs 17.20 margin.
Registered 2026-10-02 · resolved 2026-10-02
Division-aware prior (decay_od_div) fixes FCS under-rating fix works
Question. Adding unpenalized FCS offense/defense offsets (frozen decay_od hyperparameters, no retune) improves margin accuracy.
Result. HOLDOUT (2,028): 16.61 -> 16.04 (+0.58); HOLDOUT-2: +0.43. FBS-vs-FCS margin bias +12.2 -> -0.2. 2026 wk1-4 FCS-game bias +18.8 -> +3.5, which matches the closing line's own +3.5 on those games.
Verdict. PASS.
Registered 2026-10-02 · resolved 2026-10-02
decay_od adds information beyond the opening line rejected
Question. Forecast-encompassing actual = a + b*open + c*decay_od has c > 0.
Result. TEST (938 games with an opener): margin c = +0.031 (p 0.75), total c = +0.008 (p 0.96). VALIDATE margin c = -0.141 (p 0.04, wrong sign).
Verdict. FAIL. Encompassed, the same as every model before it.
Registered 2026-10-02 · resolved 2026-10-02
Recency weighting (season discount / in-season half-life) improves decay_od recency helps accuracy (season boundary only)
Question. Tuned decay_od beats the identical model with no recency (H = inf, s = 1).
Result. TUNE selected H = inf (NO in-season decay) for both markets, s = 0.5 margin / 0.2 total. Every finite half-life 7-120 days was worse, monotonically. TEST: +0.30 margin, +0.53 total; VALIDATE +0.19 / +0.56.
Verdict. PASS, with a specific shape. Discount last season's games; do not decay within the current season. This is about accuracy only; it says nothing about edge (see decay-od-encompassing).
Registered 2026-10-02 · resolved 2026-10-02
decay_od (opponent-adjusted, recency-weighted ratings) beats elo rejected
Question. decay_od, tuned on 2014-2021 (H, s, lambda grid), beats the repo's EloModel on paired residual sd against actual results.
Result. TEST 2025 (1,596 games): margin decay_od 16.78 vs elo 16.54 (-0.23), total 15.63 vs 15.72 (+0.08). VALIDATE: -0.28 margin, +0.10 total.
Verdict. FAIL. Better on totals than elo but well below the threshold, and worse on margin.
Registered 2026-10-02 · resolved 2026-10-02
Do our models add information beyond the opening market? no edge
Question. The mandate's central question: does the football model contain information about outcomes that is not already incorporated into the opening market? Tested by forecast-encompassing regression, actual = a + b*open + c*model. If c is indistinguishable from zero, the model is encompassed and contributes nothing beyond the opening line whatever its standalone accuracy.
Result. Not one model coefficient is significant, on either market, in either period (elo margin -0.076 p=0.22 / +0.065 p=0.44; elo total -0.051 p=0.68 / -0.241 p=0.19; featured margin -0.068 p=0.58 / -0.009 p=0.96; featured total +0.021 p=0.86 / -0.051 p=0.81). Several are negative. The open coefficient runs +0.86 to +1.03 with t from 6 to 34. Residual regressions return R-squared of 0.0000 to 0.0024 everywhere. One TRAIN candidate -- featured totals predicting subsequent market movement, +0.0589, t=+4.38, p<0.0001 -- flipped sign in VALIDATE (-0.0172, t=-0.95, p=0.34) and died.
Verdict. Every model is encompassed by the opening market, on both markets, in both periods. The single strong TRAIN cell did not replicate; 24 cells were inspected and significant coefficients appeared in BOTH directions, which is the signature of noise-mining rather than signal. This is the project's headline finding and the reason config.EDGE_VALIDATED is False.
Registered 2026-09-19 · resolved 2026-09-19
NIL/transfer-portal era regime change rejected
Question. The NIL/transfer-portal era (roughly 2021+) changed roster construction enough that older seasons are actively misleading. A feature model trained only on the 4 most recent seasons predicts better out-of-sample than the same model trained on all history.
Result. Paired on 1,566 identical games: featured 21.16 margin / 16.45 total; featured_recent4 21.12 / 16.46. Improvement +0.04 margin, -0.01 total.
Verdict. FAIL, and not marginally -- the difference is an order of magnitude below the registered threshold on ample sample. It does NOT mean recency is unimportant; it means the regression's training window is the wrong place to express it, because recency already lives inside the features (a 4-game half-life on in-season form).
Registered 2026-09-19 · resolved 2026-09-19
SP+ totals disagreement signal rejected
Question. The SP+ walk-forward model's TOTALS predictions beat the closing total when the two disagree materially -- taking the model's side of a 6+ point disagreement wins more than the -110 break-even rate.
Result. 49.2% on n=1502 at the registered 6.0+ threshold -- below the criterion, below a coin flip, with three times the required sample. Flat-to-declining across thresholds (50.5% / 49.3% / 49.2% / 50.5%), which is what a no-signal series looks like. Model totals sd 18.29 vs the closing line's 16.73 on the same set.
Verdict. FAIL, unambiguously. A genuine rejection, not an underpowered inconclusive. The 2022-2024 figures were noise across 8 simultaneously-inspected cells, exactly as suspected at registration.
Registered 2026-09-18 · resolved 2026-09-18