Six months and 40,000 resolved claims later: the sit edge narrowed from 12 points to 10, start calls corrected upward from noise to a small real edge, and every show is now measured against its own discussed-population baseline.
Fourteen shows, three seasons, pre-registered before results: 2024's rankings didn't predict 2025's (rho = 0.23), but sit calls beat start calls by about 13 points, every season.
We locked the test in public two weeks ago. Here's what came back — including the part that didn't confirm what we wanted.
Before we run the numbers on whether accurate shows stay accurate, here's the test, locked in public.
We rebuilt our scoring bar around startable weeks, regraded the whole corpus, and tested the results against two skill-less baselines.