Within historical rich-verdict cohorts (S&P membership at five formation vintages, 2016–2021, point-in-time filed data only), names whose durability the filings corroborate beat the hype-flagged names by a median +21.7pp over 3 years and +36.1pp over 5 years of forward excess return — positive in 5 of 5 vintages, every bootstrap confidence interval excluding zero, and the spread survives removing the heavy-capex sectors. The engine’s REFUSED cohort was the worst-performing cohort in the study (median −22.5pp / −44.6pp excess): the habit of refusing degenerate bases is forward-return-validated, not just epistemically tidy.
Replayed point-in-time (only facts filed by each decision date), the hype meter’s exit signal called 5 of 5 peak exits in the 2021–22 unwind, adjudicated against filed outcomes. The one false positive (a secular name flagged as cyclical hype) was resolved by the cycle-vs-secular companion lane — and trust-gating supplied the precision: signals fire only on valid instruments.
Extending the machinery to a 200-name $1–10B-revenue universe built from filed data only: the pipeline survives down-cap almost unchanged (data-gate rate 5.0% vs 4.2% on the S&P), and the mid-cap pond prices FAIR (median price/value ratio 1.02x) where the S&P prices rich (1.33x) — consistent with the thesis that any information edge lives where analyst coverage is thin, with fatter tails on both sides.
The standing falsification loop. The engine claims the median S&P name is rich and that its refusals deserve refusing. This scoreboard freezes mechanical portfolios monthly and scores them quarterly, forever — the claim becomes accumulating, dated data. Report-only: the scoreboard judges the engine and never feeds it.
Regenerated 2026-08-01 by scripts/scoreboard.py report.
| sleeve | rule |
|---|---|
| MECH | construct_starter, DEFAULT knobs, no owner judgment — the mechanical proposal as-is; weights as proposed, cash scored at 0% price return. top 6 factors by sleeve conviction, water-filled proportional to conviction under a 20% per-factor cap over 90% (cash floored at 10%); <= 2 names per factor, equal-weight within a sleeve |
| CHEAP | trust A/B, engine margin > +15%, equal-weight |
| RICH | trust A/B, engine margin < -15%, equal-weight — the expensive-verdict cohort (the duration question's test subjects) |
| REFUSED | trust D + engine refusal lanes (error records), equal-weight — the engine's own "do not trust" set; it should underperform |
| ALL | every artifact name with a usable price, equal-weight — the INTERNAL comparator (see comparator note); not the cap-weighted index |
| vintage | artifact | engine | MECH n | CHEAP n | RICH n | REFUSED n | ALL n |
|---|---|---|---|---|---|---|---|
| 202607 | sp500_analysis_20260731_153015.json | 4.24.1 | 11 | 51 | 132 | 148 | 503 |
No freeze is a quarter old yet — the first scores land at vintage + 3 months. The table below will accumulate forever; young samples judge nothing.
Data: data/scoreboard/freeze_YYYYMM.json (immutable per month), data/scoreboard/scores.json (append-only). This page is a regenerated view.
Every valuation is built from SEC EDGAR filings only — market prices never enter the model; they appear only in comparisons like the studies above. Full methodology: How the engine values companies →