Millions of fantasy football recommendations are made every season. Very few are ever measured.
ScoutRank exists to measure what actually happened.
We don't predict football. We measure credibility.
Every rating is built from verified calls and real outcomes — not reputation, not follower count, not opinion. Every rating can be independently verified by reviewing the underlying receipts.
A scoreable start/sit call requires (1) a specific player, (2) a specific week, and (3) a lineup-actionable stance — a statement a listener could set their lineup on.
Analysts rarely say the literal words "start him." They say "I like him this week," "he's the play," "I'd go with him." Those are calls, and we score them.
We capture (not yet scored):
A call enters the scored record only when it clearly resolves to a start-or-sit stance for a specific week. A credibility rating is only as trustworthy as the calls underneath it — so we keep the scored scope narrow on purpose, and show everything else honestly as captured-but-unscored rather than counting calls a publisher didn't clearly make. Scoring the rest is on our roadmap.
Every rating is backed by receipts: the individual calls that produced it, each showing the player, the week, the recommendation, the actual points, and the result — and the source: the exact words as spoken, the show, the air date, and a link to the episode audio.
Nothing is hidden behind the score. If you want to know why a publisher is rated the way they are, the calls are right there to check — down to the tape. This is the whole point: a rating you can audit, call by call, quote by quote.
Each rating follows the same chain, from what was said to the number:
Every scored call is checked against the player's actual half-PPR fantasy points for that week, against a bar set for his position — and that bar is recomputed fresh every week from that week's real results, not a fixed number that stays still all season:
Concretely: the bar for each position, each week, is the half-PPR score of the Nth-ranked player at that position across the full NFL player population that week (not just the players podcasts happen to discuss) — the 12th-best QB or TE, the 24th-best RB or WR. Whoever cleared that line that week was startable; whoever didn't, wasn't. A representative sample:
| Week | QB line | RB line | WR line | TE line |
|---|---|---|---|---|
| 1 | 22.6 | 10.7 | 10.5 | 9.4 |
| 9 | 22.2 | 9.0 | 12.1 | 10.0 |
| 18 | 17.3 | 8.3 | 8.5 | 8.8 |
The line moves — week 1's QB bar (22.6) is meaningfully higher than week 18's (17.3), because scoring environments shift over a season. A flat, fixed number can't track that; a bar recomputed from that week's actual results does. Every claim's receipt shows the exact line that applied to it, not a season-average approximation.
The check that keeps it honest: under these per-week lines, a skill-less caller — one making random start/sit calls on real players in real weeks — scores almost exactly 50%. We simulate this directly: 200 trials, mean 49.63%, tightly banded around the coin-flip line. That's what makes "50 = coin flip" a validated claim rather than an assumption: the only way to score above 50 is to beat random, against the same bar a real fantasy manager is judged by.
Voids. Some calls can't be scored — the player didn't play, there's no game data, the position isn't covered by a startable-lineup slot (e.g. kicker), or the recommendation wasn't a clean start/sit. These are marked void and excluded from accuracy. A void isn't a hit or a miss; it's a call that couldn't be verified, and it never counts for or against a publisher.
“Verified call” — the definition. Every count on this site labeled a “verified call” (the site-wide total, each publisher's record, every article) means the same thing: a claim with a hit or miss outcome. Voided claims are excluded from that count entirely — not scored as a loss, not scored at all. This is the same figure the leaderboard's own “verified calls” total is built from, so the number you see on the board and the number cited in any ScoutRank analysis will always match.
ScoutRank combines three things:
Why not just use raw accuracy?
Because accuracy alone is misleading on small samples. Consider two publishers:
| Publisher | Record | Raw Accuracy |
|---|---|---|
| Publisher A | 18 of 24 | 75% |
| Publisher B | 321 of 500 | 64% |
Raw accuracy says Publisher A is better. ScoutRank says Publisher B — because 24 predictions don't prove as much as 500. A great record on a handful of calls hasn't yet earned the same trust as a strong record across hundreds.
This is the correction that makes that possible. ScoutRank pulls small-sample ratings toward the middle until enough evidence accumulates. A high accuracy on a few calls produces a strong raw number but a more cautious rating — because the rating reflects what's been proven, not just what's been observed. (Statistically, it's a shrinkage adjustment; in plain terms, it's evidence.)
ScoutRank is normalized to 0–100, where 50 is a coin flip — the line between better and worse than chance. Higher is better. Because real start/sit accuracy clusters in a relatively narrow band, the ratings are intentionally compressed: the differences between publishers are real, but not dramatic, and we don't inflate the spread to make them look bigger than they are.
Every rating shows a confidence level alongside it — Provisional, Low, Moderate, High, or Very High — reflecting how much verified evidence stands behind it. A high ScoutRank with low confidence means "promising, but not yet proven." Confidence never changes the rating itself; it tells you how much to trust its stability.
A full rating requires a full record. Publishers whose scored calls are concentrated in a single position carry a limited position coverage note and are listed with the developing publishers rather than the main board — excellence at one position is real, but it isn't yet a complete credibility record. This is a statement about evidence breadth, not quality.
Each publisher's record is also shown split by direction — accuracy on start calls and on sit calls separately. The two are genuinely different skills, and across the whole corpus, publishers are markedly better at start calls than sit calls. The split is shown so you can see which half of a record does the work.
A credibility rating should be transparent about its own limits.
Listing is not optional — public predictions get public review, and no publisher can be removed from the record. Publishers may opt out of promotional use: on request, we won't feature their show in our marketing, social posts, or highlighted content. Their scores, receipts, and profile remain public and identical to every other publisher's.
Publishers may elect anonymized display on ScoutRank. On request, the publisher's name, logo, and outbound links are replaced with a stable anonymous label on all public-facing display.
Electing anonymized display changes nothing about the research record. All underlying claim data, timestamps, direction labels, and resolved outcomes remain permanently in the corpus and in all aggregate analyses. No claims are removed, voided, or altered. The publisher's accuracy and ranking are still computed from the same data under the same methodology.
This policy is in beta. The standard display — real name, full profile, public receipts — remains the default. Display-layer anonymization standardizes with SR-2027 before the next season begins.
ScoutRank evolves through versioned methodology updates. The current version is SR-2026.4.
When we improve it, we publish a new version and document what changed. We don't silently rewrite past ratings — every recomputation is logged, and prior versions remain auditable.
The scoring bar is frozen for the 2026 season — corrections to data will continue; the bar itself will not change again until season's end.
2025 season backfill in progress — publisher scores update as full-season coverage completes per show.
An earlier version of the Late-Round Fantasy Football audit page stated that graded claims came only from mainline start/sit episodes. The underlying scan had no category for sleeper-format episodes at all, so every "Week N Sleepers" title fell through unclassified. A follow-up audit found sleeper-format episodes account for a majority of this publisher's graded starts, graded under the same rules applied to every show. The audit page was corrected with a permanent notice; the scan was fixed.
A previously-held content source briefly entered the extraction pipeline due to a code gap: FantasyPros publishes its football show through a dedicated feed, but a broader multi-show feed under the same publisher name was meant to stay excluded pending review and wasn't. We caught it, quarantined the affected claims immediately, verified them for genre accuracy (100% genuine start/sit fantasy football content, no cross-sport bleed) and for duplication against the publisher's existing record (no meaningful overlap found), then reinstated and scored them normally. FantasyPros' rank moved from 9 to 10 as a result — more verified data, not a ratings adjustment. The broader feed remains excluded from future extraction.
Four publishers' records were split across duplicate identities on the leaderboard — the duplicates weren't hidden, they showed under the same name at real ranks (one at rank 2). A platform-matching quirk let a second identity get created for a show that already had one, so its most recent episodes attributed to the wrong id while the original sat on stale, months-old data. Merged 2026-07-28: episodes double-counted under both ids were resolved (the older extraction voided, not deleted — its permalink still resolves honestly), the rest reassigned to the real identity, and the duplicate rows removed. Two of the four publishers' combined scores came out worse than their duplicate-alone scores had shown — the merge is honest, not a ratings boost.
Added a promotional opt-out policy — publishers may request exclusion from Scout's marketing, social posts, and highlighted content while remaining fully listed with public scores, receipts, and profile. Applied to one publisher on request.
Routine claim-validity audit of a newly ingested publisher; 10 claims voided for scope/genre misclassification (season-long, retrospective, or availability takes misread as weekly start/sit stances), 1 corrected for a week-assignment error. Full methodology unchanged.
SR-2026.2 → SR-2026.3 — Corpus Integrity Release. Before the 2026 season, we audited our entire corpus — every claim, every quote, every verdict — against the source recordings. We found and fixed four defects: duplicate episode ingestion that inflated claim counts; a small share of trade and ranking statements mislabeled as start/sit calls; claims whose quoted evidence referenced a different player than the one scored; and claims that weren't falsifiable lineup calls at all (rankings without a decision, DFS pricing, betting props, general praise). Every affected claim was voided or corrected — voids never count for or against a publisher. Net effect: rankings essentially unchanged. The #1 publisher didn't move, no publisher moved more than a few positions, and most accuracy figures ticked slightly up, because what we removed was noise, not signal. The verified-call total decreased from 2,610 to 2,404 — the corpus got smaller and truer. We also added two publishers with full 2025 records: The Audible (Footballguys) and the Fantasy Life Show. Every quote on ScoutRank is now verified word-for-word against the source recording before publication.
SR-2026.1 → SR-2026.2. The original methodology used a single flat 10-point bar for every position. Community review after launch identified what our own audit confirmed: the flat bar was position-biased — near-automatic for quarterback starts, structurally harsh on tight ends — which meant ratings partly reflected position mix rather than skill. SR-2026.2 replaced it with position-calibrated median thresholds, validated by naive-caller simulation, and added the position-coverage rule and the start/sit split. Some ratings changed materially under the correction, including findings we had featured prominently. We published the changes ourselves, in full. That is what the versioning is for.
Future versions may include scoring for rankings, waivers, and trades — deepening the record from start/sit accuracy toward a fuller picture of publisher credibility.
Some questions are still open. We have not decided them yet. Any change we make here will apply to every show at once. It will happen after a season ends, not in the middle of one. We will never rewrite past results.
Should speculative segments be scored in a separate tier? Some shows run segments built around long-shot picks: sleepers, deep stashes, low-confidence dart throws. Today we score these the same way we score every other start/sit call. Whether they belong in a separate tier is an open question. If we make this change, it applies to every show on the site, not one. It takes effect after a season ends, and it does not touch calls already scored.
How do we decide what a call means? Analysts rarely say the literal words "start him" or "sit him." We read the analyst's stated stance in context: what they actually said about a specific player for a specific week. Every graded claim links to the source audio, so you can hear the exact words and judge whether we read the stance correctly.
On their Week — (—) episode, The Fantasy Footballers explicitly recommended starting —. The recommendation is captured verbatim and timestamped before kickoff.
ScoutRank is independent. We rate predictions, not personalities.