Within historical proxy-rich cohorts (S&P membership at five formation vintages, 2016–2021, point-in-time filed data only), names whose durability the filings corroborate beat the hype-flagged names by a median +21.7pp over 3 years and +36.1pp over 5 years of forward excess return — positive in 5 of 5 3y vintages. The descriptive pooled bootstrap intervals exclude zero, but overlap across issuers and time windows makes them optimistic; the spread stays positive after removing heavy-capex sectors. The proxy-refused cohort was the worst-performing cohort in the study (median −22.5pp / −44.6pp excess): these proxy results do not validate the engine’s refusal rules or demonstrate prospective investment usefulness. Method: 2026-07-31 Yahoo adjusted-close ratios, including dividend and split adjustments; approximate total returns. SPY adjusted-close return over the identical window; excess is in percentage points. Five formation vintages: 2016, 2017, 2018, 2019 and 2021; pooled durable-rich n=201, hype-rich n=269, proxy-refused n=223 (name-vintage observations). Historical S&P membership; financial SIC 6000–6799 excluded; unmapped/acquired/delisted names excluded from classified cohorts. Point-in-time filed facts classified by a fixed-anchor CFO-minus-capex proxy, not an engine replay. 3y and 5y returns; the 2021 5y window is partial through 2026-07-30. Limits: Survivorship: excluded names were not classified. The direction of bias in the durable-rich minus hype-rich spread is unknown; acquisition outcomes are also missing. Pooled bootstrap intervals are descriptive and optimistic: repeated issuers and overlapping time windows are not independent observations. Five positive 3y vintages are not five independent replications. The proxy-refused cohort is not the engine REFUSED cohort. These results do not validate engine refusal rules or demonstrate prospective investment usefulness.
Replayed point-in-time (only facts filed by each decision date), the hype meter’s exit signal called 5 of 5 peak exits in the 2021–22 unwind, adjudicated against filed outcomes. The one false positive (a secular name flagged as cyclical hype) was resolved by the cycle-vs-secular companion lane — and trust-gating supplied the precision: signals fire only on valid instruments. Method: Selected point-in-time exit-signal cases adjudicated against the 2021–22 unwind. No aggregate return statistic reported in this receipt. No investment-return benchmark reported in this receipt. Limits: Five selected historical cases do not estimate prospective accuracy or investment returns.
Extending the machinery to a 200-name $1–10B-revenue universe built from filed data only: the pipeline survives down-cap almost unchanged (data-gate rate 5.0% vs 4.2% on the S&P), and the mid-cap pond prices FAIR (median price/value ratio 1.02x) where the S&P prices rich (1.33x). These cross-sectional ratios do not establish an information edge or better future returns. Method: 2026-07-31 200-name mid-cap pilot from filed $1–10B-revenue companies. Cross-sectional engine price/value ratios and data-gate rates. No forward-return statistic: these are valuation ratios, not investment outcomes. Contemporaneous S&P engine cohort; not a return benchmark. Limits: A lower model valuation ratio does not establish an information edge or better future returns.
The standing falsification loop. The engine claims the median S&P name is rich and that its refusals deserve refusing. This scoreboard freezes mechanical portfolios monthly and scores them quarterly, forever — the claim becomes accumulating, dated data. Report-only: the scoreboard judges the engine and never feeds it.
Regenerated 2026-08-01 by scripts/scoreboard.py report.
| sleeve | rule |
|---|---|
| MECH | construct_starter, DEFAULT knobs, no owner judgment — the mechanical proposal as-is; weights as proposed, cash scored at 0% price return. top 6 factors by sleeve conviction, water-filled proportional to conviction under a 20% per-factor cap over 90% (cash floored at 10%); <= 2 names per factor, equal-weight within a sleeve |
| CHEAP | trust A/B, engine margin > +15%, equal-weight |
| RICH | trust A/B, engine margin < -15%, equal-weight — the expensive-verdict cohort (the duration question's test subjects) |
| REFUSED | trust D + engine refusal lanes (error records), equal-weight — the engine's own "do not trust" set; it should underperform |
| ALL | every artifact name with a usable price, equal-weight — the INTERNAL comparator (see comparator note); not the cap-weighted index |
| vintage | artifact | engine | MECH n | CHEAP n | RICH n | REFUSED n | ALL n |
|---|---|---|---|---|---|---|---|
| 202607 | sp500_analysis_20260731_153015.json | 4.24.1 | 11 | 51 | 132 | 148 | 503 |
No freeze is a quarter old yet — the first scores land at vintage + 3 months. The table below will accumulate forever; young samples judge nothing.
Data: data/scoreboard/freeze_YYYYMM.json (immutable per month), data/scoreboard/scores.json (append-only). This page is a regenerated view.
Caveat corrected 2026-09-14: the prior text asserted that missing outcomes flattered the affected sleeves, particularly REFUSED. That direction was not measured. Counts, returns and the original report date above are unchanged.
Company fundamentals come from SEC EDGAR filings; valuation also uses macro inputs and explicit modeling assumptions. Market prices appear in comparisons, not intrinsic-value calculations. Full methodology: How the engine values companies →