We Value Your Privacy

    We use cookies to enhance your experience, provide essential functionality, and analyze site usage. You can customize your preferences or accept all cookies. Learn more

    Strategy Library
    The kill log

    A vetting bar only means something if you can see what it rejects.

    Start with the question you came for

    Below: the full archive, sweep by sweep, in the order it was run. Six strategies cleared this bar — read the ones that survived. The bar itself gained a sixth check in September 2026, after a candidate passed the other five and turned out to be indistinguishable from random entry timing.

    Research Log — what was tested and killed

    The library only ships survivors, which makes it easy to forget how many candidates die. This log records completed research sweeps: the pre-registered hypotheses, the protocol, and the full results — including every kill. Two reasons it exists:

    1. Honesty. "Why so few entries?" Because this is what the vetting bar does to ideas.
    2. Discipline. A documented kill must not be quietly re-mined later with new knobs until it looks good. If a dead idea is to be revisited, it needs a genuinely new mechanism or market structure change — stated up front.

    Sweep 2026-08-05 — lower timeframes (crypto 1h/4h)

    Question: is there a strategy on lower timeframes (1h primary; 4h probe; 15m if anything survived 1h) that clears the vetting bar?

    Protocol. Pre-registered hypotheses and neighbors (fixed before any run; neighbors run only for families whose primary showed genuine signal). Engine: the app's own (runBuilderBacktest), long entries at bar close. Costs: 0.1%/side fees — corrected from sweep #4 on to the USDT-M perpetual schedule, 0.05% taker / 0.02% maker — plus perp funding, charged at each 8-hour settlement a position is held through (not per bar: a 15m bar that crosses no settlement pays nothing). Assets: BTC, ETH, BNB (majors). Two independent windows: W1 Aug 2018 → Aug 2022, W2 Aug 2022 → Aug 2026. Funding history begins Sep 2019, so funding-gated setups effectively start there in W1 — and holds before that date are charged no funding at all. Gates per asset: net positive in BOTH windows, survives strip-best-trade in both, max DD ≤ 20%, then cross-asset (≥ 2 majors).

    What the three headline numbers actually measure. Return is the arithmetic sum of per-trade percentages, not a compounded equity curve — which is why dividing it by the trade count is a legitimate per-trade average, and why +887% is not a 9× account. Drawdown is measured on that same summed curve and marked only at trade exits, so a position that goes far underwater and closes green contributes nothing to it; it is quoted in points, and where the two differ the account-basis figure follows in parentheses. MC p95 is a shuffle of the same trades into a different order, repeated 1,000 times — it bounds how unlucky the sequencing could have been, and says nothing about return uncertainty. All three are gated on the stricter reading where a choice exists.

    Hypotheses and verdicts

    # Hypothesis (pre-registered spec) TF Verdict
    H1 Donchian breakout long: close crosses above highest(prior 55), ATR-trail 2×ATR(14). Neighbors: period 20/100, trail 3× 1h, 4h Killed — see below; nearest miss of the sweep
    H2 Squeeze-breakout long: Donchian-48 width ≤ 3% AND close crosses above highest(48), ATR-trail 2× 1h, 4h Killed — window-unstable
    H3 RSI(2) < 10 above SMA(200), exit close > SMA(5) (entry #2 verbatim on crypto 1h) 1h Killed — decisively
    H1s Short mirror: close crosses below lowest(55) 1h Killed — decisively
    H4 Entry #1 funding-squeeze verbatim at 1h (funding < 0 AND RSI(14) crosses above 30, ATR-trail 2×) 1h Killed as a library entry — BTC-only and marginal

    H1 — Donchian breakout (the nearest miss)

    On 1h: destroyed everywhere (PF 0.74–0.92, −31% to −102% per window, 300–500 trades). Fees and chop churn dominate. The 20-period neighbor is even worse; 100 doesn't save it.

    On 4h, the family shows a real momentum edge — with the pre-registered 3×ATR trail (dc55x3) all six asset × window cells are net positive and all six survive strip-best-trade:

    Asset W1 W2
    BTC +87.4% · PF 1.52 · DD 31.4% +37.7% · PF 1.35 · DD 32.7%
    ETH +124.9% · PF 1.50 · DD 75.9% +36.2% · PF 1.23 · DD 27.2% (strip-best +0.4%)
    BNB +173.6% · PF 1.88 · DD 45.3% +59.3% · PF 1.53 · DD 26.1%

    Why it's not an entry: drawdown. 26–76% max DD versus the bar's ≤ 20%. This is structural for an always-in breakout system on crypto — ~35% winners with long losing streaks through chop. The edge is real; the raw form is not compatible with a sane risk plan, and dressing it up (extra filters until the DD looks nice) is exactly the curve-fit the bar forbids. If it ever ships it will be as an explicitly high-drawdown, sleeve-sized, paper-trade-first entry — a decision to make deliberately, not by quietly relaxing the bar.

    H2 — squeeze-breakout: the two-window gate doing its job

    BTC W1 looked spectacular (58 trades, PF 3.73, +51.8%, DD 6.3%, survives strip-best) — and W2 collapsed (168 trades, PF 0.53, −37.3%). ETH roughly mirror-image, BNB small both ways. Same spec, adjacent windows, opposite result: this is what an unstable regime artifact looks like, and publishing the W1 row alone would have been a lie of omission. Neighbors were not run (family failed its primary — pre-registration rule). On 4h the setup never occurs (an 8-day ≤ 3% range doesn't happen on crypto).

    H3 — mean reversion does not transplant to crypto 1h

    Entry #2's exact spec, which is genuinely robust on daily index ETFs, loses catastrophically on crypto 1h: PF 0.51–1.19, −40% to −140% per window, ~750 trades each. A diversified index mean-reverts on a daily panic; a single crypto asset on a 1h dip is just as likely trending to zero for the day. No-stop dip-buying is knife-catching here.

    H1s — shorts still have no edge

    PF 0.56–0.88, −46% to −166% everywhere. Consistent with entry #1's finding: crypto's drift and squeeze dynamics make fading weakness a losing short. Long-only stands.

    H4 — funding-squeeze at 1h does not generalize

    Entry #1's write-up reports a flattering 1h row (BTC recent-24mo: PF 1.67, survives strip-best). Across the full history and assets: BTC 12.6%/6.5% per window but PF ~1.2, a 20.5% DD breach, and W2 strip-best of +0.2%; ETH negative both windows; BNB +74%/−12.5%. BTC-only and marginal ≠ library-grade. The 4h original remains the tradable form; the entry #1 doc now carries this cross-window note on its 1h row.

    Structural conclusion

    At realistic retail costs (0.1%/side + funding), no 1h crypto strategy expressible in the current strategy DSL cleared the bar — momentum churns, mean reversion catches knives, squeezes don't repeat across regimes, shorts lose. The per-trade edge needed to overcome ~0.2% round-trip plus noise simply isn't there at this frequency for these mechanism families. 15m was not run: it strictly worsens the cost-to-edge ratio that already killed 1h. Lower-timeframe research that wants a different answer needs a different lever: maker/low-fee execution assumptions (a different product than this library models), intraday session/time-of-day structure (not currently compilable), or order-flow data beyond OHLCV + funding.

    Superseded in part by sweep #4 (below). This conclusion charged 0.1%/side, the spot-taker default; these instruments are USDT-M perpetuals at 0.05% taker. Re-run at the corrected rate, the 1h funding-squeeze passes on BTC and ships as entry #5, and the 1h mean-reversion result turns out to have been gross-POSITIVE at roughly +95% — the execution model killed it, not the signal. "Mean reversion catches knives" is not the standing verdict; "regime-fragile, not cost-dead" is. The structural claim that survives is narrower: at real perp costs the 1h layer is reachable, but the edges there that do not read market internals appear bound to the volatility era they were born in.


    Sweep 2026-08-06 — session structure (crypto 15m/1h, new session atoms)

    The previous sweep's "different lever": session/time-of-day atoms were added to the strategy DSL (hour_utc, day_of_week, session open/high/low/VWAP, opening ranges — fixed UTC windows), unlocking the classic intraday literature. Same protocol, assets, windows, costs, and gates as the 2026-08-05 sweep; 15m primary, 1h probe.

    Hypotheses and verdicts

    # Hypothesis (pre-registered spec) Verdict
    S1 US-open ORB long: first hour of the US session {13.5–20 UTC}, close crosses above the range high (no entries after 18:00), 2×ATR trail + session-end exit Killed at costs — see the gross diagnostic below
    S1s Short mirror (one probe) Killed — −135% to −254% per window
    S2 Session-VWAP reversion long: close < US-session VWAP × 0.99, exit on VWAP reclaim or session end Killed — negative in all 12 cells
    S3 London breakout of the Asia range {00–07 UTC}, entries 07–13:30, 2×ATR trail + session-end exit Killed — W1/W2 regime flip, even gross

    All 42 net-of-cost cells were negative (PF 0.38–0.98). No hypothesis reached the neighbor or stress phases.

    The gross diagnostic — quantifying WHY (fees off, funding still on; not tradable)

    Setup (gross) W1 W2 Reading
    ORB BTC 15m +28.8% · PF 1.09 · sb+ +36.3% · PF 1.17 · sb+ Real, window-stable gross tendency
    ORB ETH 15m +31.5% · PF 1.08 · sb+ +43.0% · PF 1.16 · sb+ Same
    VWAP-reversion (best cells) mixed + mostly − Unstable even gross
    London breakout big + in W1 negative everywhere in W2 Regime artifact even gross

    The US-open ORB on BTC/ETH 15m is the interesting kill: a genuine session-structure edge exists — positive in both independent windows, survives strip-best, about 870 trades per window. But it averages about +0.04% per trade gross, and break-even for an edge that size is 0.02%/side.

    This sweep charged 0.1%/side (spot-taker default), which puts the edge 5× below the cost line. Sweep #4 corrects the instrument: at the USDT-M perpetual taker rate of 0.05%/side the round trip is 0.1% and the edge is 2.5× below it; at the 0.02% maker rate it is break-even. The honest retail number is 2.5×. Break-even sitting on the maker rate is not a route in for this family — a breakout enters by crossing a level, which is a taker action by construction. That is the low-timeframe conclusion in one sentence: the structure is there, and retail execution costs are bigger than it is.

    Session atoms remain in the DSL — they are useful as filters and exits on higher-timeframe strategies (session-end exits, day-of-week gates) even though no standalone intraday entry cleared the bar.


    Sweep 2026-08-06 (#3) — new instruments & mechanisms, daily bars

    Pre-registered: (T1) entry #2's RSI(2) dip-buy verbatim, zero knobs on sector SPDRs (XLK/XLF/XLE/XLV/XLI) + international (EFA/EEM); (T2) the 2026-08-01 pre-registered leftover, a crypto daily trend-pullback (close > SMA200 AND RSI(14) < 40, 3×ATR trail, BTC/ETH/BNB); (T3) SMA-200 trend timing (Faber-style cross, SPY/QQQ/EFA). Same gates; new evidence attached from the holistic engine: DD-from-peak, exposure, Monte-Carlo p95 shuffled drawdown, 0.1%/side fee stress, and next-open fill checks.

    Hypothesis Verdict
    T1 sectors/international 5 of 7 PASS → shipped as entry #2's extension section (XLK/EEM first-rank; XLF/XLE/EFA baseline-costs-only; XLV failed W2, XLI breached the DD bar by 0.1)
    T2 crypto trend-pullback Killed — window-flipping everywhere (BTC/ETH: W1 negative, W2 positive; BNB the mirror). The leftover is closed; no neighbors run
    T3 SMA-200 timing SPY/QQQ PASS → shipped as entry #3. EFA failed (−4.3% W1). The 2000–2010 stress decade fails strip-best — reported in the entry as the known whipsaw cost

    Also measured: the dip-buy family is genuinely flattered by close fills (SPY W2 +36.4% → +19.4% at next-open — still profitable, now quantified in entry #2); the trend-timing system is fill-indifferent (next-open slightly better). Costs kill nothing at ~3 trades/yr; they thin the dip-buy's secondary instruments exactly as they did IWM.


    Sweep 2026-08-06 (#4) — low timeframes under corrected market structure

    The sweep-1/2 kill rule allows a revisit only with a stated market-structure change. Two were stated up front: (a) cost correction — prior sweeps priced crypto at 0.1%/side, the spot-taker default, but these strategies trade USDT-M perpetuals whose standard schedule is 0.05% taker / 0.02% maker; (b) new engine capability — resting limit-order entries with strict-pierce fills at the maker rate (mean-reversion entries are naturally passive). Motivating arithmetic, stated before running: sweep-1's 1h RSI(2) MR lost about 55%/window net while paying about 150% of costs — it was gross-POSITIVE at roughly +95%; the execution model, not the signal, killed it.

    # Hypothesis Verdict
    P1 Entry #1's funding-squeeze at 1h, real futures taker fees (spec verbatim) BTC passes every per-asset gate (W1 +18.2% PF 1.39 DD 17.0 sb+; W2 +13.1% PF 1.36 DD 11.6 sb+; 122 trades). ETH/BNB still fail → BTC-only → ships as entry #1's rehabilitated 1h variant, not a standalone entry
    P2 Passive limit MR, 1h (limit at close −0.5%, maker 0.02 in / taker 0.05 out) Killed — but the thesis was right in W1: BNB +179.5% PF 1.69, ETH +112.1%, BTC +59.0%, all strip-best-surviving — passive fills DO capture the gross edge. Then W2 flips BTC/ETH negative and BNB breaches the DD bar in W1. Regime-fragile, not cost-dead. No neighbors run (no full passer)
    P3 Sweep-1's Donchian-4h 3×ATR + a daily SMA-200 trend filter (the single pre-registered revisit variant) Killed — the filter worked as theorized (drawdowns roughly halved; BTC W2 now 18.4, under the bar) but W1 drawdowns still breach (30–64). All 6 cells positive and strip-best-surviving, again. The revisit clause for this family is now SPENT

    Structural conclusion, updated

    The corrected market structure moves the line but doesn't erase it. With real futures fees and honest passive fills, the 1h layer contains exactly one durable edge — BTC funding-squeeze — and one regime-fragile one (passive MR: pays handsomely in the high-volatility 2018–2022 era, flat-to-negative in the post-2022 regime). The earlier "costs are bigger than the structure" conclusion softens to: at real perp costs the structure is reachable, but 1h edges that don't read market internals (funding) appear regime-bound — they need the volatility era they were born in. Anything further down (15m) remains cost-dead even at maker rates (sweep #2's ORB arithmetic).


    Sweep 2026-08-06 (#5) — metals, Double-7s, and a blocked intraday candidate

    Pre-registered: (Q1) entry #3's SMA-200 timing spec verbatim on metals (GLD/SLV, IAU as parity probe); (Q2) Connors' Double-7s channel dip-buy, frozen 2009 params (stated deviation: the DSL's rolling extremes use lows/highs where the book uses closing extremes), on SPY/QQQ/XLK/EEM; (Q3) SPY/QQQ 15m session-VWAP reversion — registered while blocked on data, run later the same day once Alpaca keys landed, spec pinned before the run (0.3% stretch below the day-anchored session VWAP, RTH SIP bars, 0.02%/side, windows 2016–2021 / 2021–2026).

    # Hypothesis Verdict
    Q1 metals trend timing Killed — GLD/IAU window-flip (2011–2015 gold bear: −18% W1, PF 0.47; then +131–135% W2). SLV positive both windows but fails strip-best in W1 (−36.2) at a 41% drawdown. Trend timing on metals is regime-hostage; the equity result does not transfer
    Q2 Double-7s QQQ + XLK PASS the full bar — including the 2000–2010 stress decade on all three US instruments (QQQ +75.7% PF 2.77 sb+). SPY is a near-miss (every check passes except W2 sum-point DD 22.6 vs ≤ 20 — covid held with no stop; account-basis DD 19.1). EEM fails on DD. Fee-stress at 0.1%/side: robust (PF 1.6–2.4). → shipped as entry #4 with a prominent correlation disclosure: same mean-reversion family as entry #2, one dip-buy sleeve, not two

    Q3 — intraday index-ETF VWAP reversion: killed, and it generalizes the low-TF finding

    Run on real regular-hours SIP bars (extended hours filtered, DST-aware), 2 bp/side — about the friendliest honest cost model retail can claim. All four cells fail: SPY nets to ~zero (PF 0.98 / 1.00, −3.1% / +0.3%, fails strip-best in both windows), QQQ is outright negative (PF 0.94 / 0.97) with 24–42-point drawdowns and Monte-Carlo p95 tails of 34–55. With ~500–1,000 trades per window paying ~0.04% round trips, the implied gross edge is ~+0.02–0.04% per trade — once again the size of the cost line, this time on the most liquid instruments in the world at near-zero fees. No neighbors were run (no passer).

    That completes the cross-asset picture the crypto sweeps started: the sub-daily mean-reversion structure exists faintly everywhere and clears trading costs nowhere — not on crypto perps at maker rates, not on index ETFs at 2 bp. The library's intraday answer stands on evidence from both asset classes now: the only durable sub-daily edge found in five sweeps reads market internals (funding), everything else lives on daily bars and above.


    Sweep 2026-08-08 (#6) — the last untested mechanism family, and an elevation

    Pre-registered: (V1) capitulation-volume reversal on crypto 1h — volume internals, the only DSL data channel never yet tested at low timeframes. Spec: volume > 3× its 20-bar average AND close below the lower Bollinger(20,2) → fade at the close; exit on a mid-band reclaim or a 2-day backstop; futures fees + funding; BTC/ETH/BNB, the standard windows and gates. Prior stated up front as moderate (Wyckoff selling-climax literature is an equities result; crypto evidence anecdotal).

    V1: killed decisively. PF 0.62–0.75 with −50% to −129% per window on BTC and ETH (BNB's lone flat cell fails strip-best). On crypto 1h, an extreme-volume flush bar is continuation, not exhaustion — the knife-catching result again, now confirmed on the volume channel too. No neighbors were run.

    Elevation, not a new find: with every price-shape, session, execution-model and volume hypothesis now tested and dead below daily bars, the sole sub-daily survivor of six sweeps — the BTC funding-squeeze at 1h on real futures fees (validated in sweep #4) — has been promoted from a footnote in entry #1 to standalone entry #5 under the README's narrow-clearer rule, with its BTC-only and fee-tier conditions stated as first-class guardrails. The library's answer to "a profitable low-timeframe crypto strategy" is that one exists, it reads market internals rather than price shapes, and everything else we could express was tested and is on this page.


    Sweep 2026-08-08 (#7) — the last rung: 15m

    One pre-registered run, no neighbors possible (the spec is frozen): the funding-squeeze — the sole sub-daily survivor of six sweeps — tested at 15m under the same futures-fee correction that validated it at 1h. Entry #1's doc had killed 15m only at spot fees on a single window; this closes it properly: full history, BTC/ETH/BNB, both windows.

    Dead. BTC — the one asset where the edge lives at 1h — is negative in both windows (PF 0.72 / 0.86, −22.9% / −10.3%); ETH negative in both; BNB's lone positive window is followed by a flat one that fails strip-best. Nothing passes anything.

    That completes the cleanest cost-vs-structure demonstration in this log: one frozen spec, three timeframes — strong and cross-asset at 4h (entry #1), BTC-only at 1h on futures fees (entry #5), gone at 15m. The per-trade edge shrinks with the bar; the per-trade cost doesn't. The sub-daily floor for this library is 1h, measured — not assumed. Below it, every mechanism family expressible on OHLCV + funding is now tested and dead; going lower is a data-acquisition problem (order flow, historical open interest, liquidation feeds — none freely available), not a backtesting problem.


    Sweep 2026-08-08 (#8) — the angle change: collect the transfers, not the candles

    After sweep #7 measured the 1h floor, the pre-registered angle change: stop predicting sub-daily candles (dead, sweeps #1–#7) and harvest the mechanism that makes them hostile — funding-rate cash-and-carry (long spot + short perp, market-neutral, collect the 8h payments). The single-leg engine cannot express two legs, so this ran on a dedicated simulator with real spot klines, real perp klines, and real funding history; rules pinned before the run (enter mean-of-3 > 0.01%/8h, exit < 0.003%; full costs both legs; returns on total capital; basis P&L from real prices).

    PASS — BTC + ETH clear every gate in both windows (worst episode across all runs: −0.2% of capital). BNB fails W2 strip-best at the primary gate but passes at the stricter pre-registered neighbor. The neighbor grid is monotone and mechanically expected: a looser entry gate fails W2 everywhere (weak funding churns the fee hurdle), a stricter one passes 6/6 — selectivity IS the edge. Shipped as entry #6 with the honest characterization stated first: while-deployed APR ~6–14%, but idle 70–90% of the time and below cash rates overall in the post-2022 quiet regime — an opportunistic harvester for hot-funding episodes, not a yield product. Manual two-leg execution (the DSL is single-leg); validated by simulator, so the entry ships without an interactive chart.


    Sweep 2026-09-18 (#9) — daily and weekly crypto: the literature transplants

    Pre-registered on 2026-09-18 and committed before the run (this section's git history is the timestamp; results are appended in a later commit). Restarting the cadence after the August sweeps: one sweep, one question, hypotheses and gates fixed up front.

    Question: does any daily-or-slower crypto strategy clear the vetting bar? Every crypto sweep so far ran at 4h and below; the only daily crypto spec ever tested is sweep #3's trend-pullback (killed). Daily and weekly bars are the untested layer — and the one where costs stop mattering, so a price-shape edge has its fairest hearing.

    What this sweep deliberately does NOT do: no sub-daily crypto (floor measured at 1h, sweep #7), no Donchian-family breakouts (revisit clause spent, sweep #4), no crypto mean-reversion dip-buys (killed at 1d in sweep #3 and at 1h in sweep #1). Three frozen published rules only, zero parameter changes, run on instruments their authors did not tune them for — the transplant discipline of sweep #3/#5.

    Protocol. Engine: the app's own (runBuilderBacktest), harness committed at scripts/research/sweep9-crypto-daily.ts (public Binance spot data, cached; reproducible with no keys). Assets: BTC, ETH, BNB spot. Windows: W1 Aug 2018 → Aug 2022, W2 Aug 2022 → Aug 2026. Costs: 0.1%/side spot fees (the library's conservative crypto default); fee stress at 0.2%/side. Fills at the bar close — crypto trades 24/7, so the next open is the close and a next-open check is a no-op. Evidence per cell: trades, profit factor, net return, strip-best, sum-point max drawdown (account-basis in parentheses), Monte-Carlo p95 shuffled drawdown, exposure, the largest open-trade giveback (highest close to exit — the number the equity curve hides on a trend system), and same-window buy-and-hold.

    Gates per asset (unchanged): net positive in BOTH windows, survives strip-best in both, sum-point max DD ≤ 20. Cross-asset: ≥ 2 majors pass, or a single-asset passer ships under the README's narrow-clearer rule with BTC-only guardrails stated first-class (the entry #5 precedent). A drawdown near-miss (≤ 25) is reported as a near-miss, not shipped.

    Neighbors run only for a family whose primary shows genuine signal — positive AND strip-best-surviving in both windows on ≥ 2 assets, drawdown ignored — and exist to show a plateau, never to pick a winner. Only T1 has neighbors (SMA-100 / SMA-300); T2 has no parameters; T3's spec is frozen.

    Hypotheses (specs fixed before the run)

    # Hypothesis Source Spec (verbatim) TF Stated prior
    T1 SMA-200 trend timing — entry #3 transplanted verbatim to crypto Faber (2007), daily twin of the 10-month SMA close crosses_above sma(200) → long; close crosses_below sma(200) → exit at close; no stop 1d Return gates likely pass (the crash years are spent in cash); drawdown is where this is expected to fail — crypto whipsaw stacks run 2–3× an equity index's
    T2 One-week time-series momentum Liu & Tsyvinski, Risks and Returns of Cryptocurrency, RFS 2021 (documented on 2011–2018 data — both windows here are out-of-sample) Weekly bars: long at the close of an up-week (close is_rising), flat at the close of the first down-week (close is_falling); no stop; no parameters 1w Moderate: the effect is published and pre-dates both windows; ~15–25 round trips a year make the 0.2% cost line a real hurdle
    T3 TTM-squeeze breakout — the 2026-08-01 pre-registered leftover ("crypto vol-squeeze breakout"), never run at 1d (the 1h cousin died window-unstable in sweep #1) Carter (2005), frozen 20 / 2.0 / 1.5 Squeeze on while bb_upper(20,2) < ema(20) + 1.5×ATR(20); long fires when the upper Bollinger crosses_above that Keltner line with close > sma(20); exit close crosses_below sma(20); no stop 1d Low–moderate: compression → expansion is a real regularity, but breakout systems on crypto have failed the drawdown gate every time

    Stated T3 deviations (the book's momentum histogram and its exit-on-decline are not DSL atoms): momentum = close above the 20-day midline; exit = close back below it. These are fixed here, before the run, and will not be adjusted afterwards.

    Also run, informational only: T1 executed on the perpetual instead of spot (0.05%/side taker, funding charged over multi-month holds) — the number a futures trader needs.

    Decision rule, written down first. A full-bar passer on ≥ 2 majors ships as entry #7. A BTC-only full-bar passer ships BTC-only with guardrails. Anything else is a kill or a documented near-miss, and the library's answer to "a profitable daily crypto strategy" becomes this section.

    Results — all three killed, and the rejection reason has moved

    Every cell's failure is a drawdown failure. Not one of the eighteen primary cells was rejected for costs. Above the daily bar the 0.1%/side fee line is irrelevant — T1 trades ~4/yr, T2 ~13/yr — and the fee stress at 0.2%/side changes no verdict anywhere. That is the sweep's structural finding: below the daily bar crypto strategies die of costs; above it they die of drawdown. The library has now measured both walls.

    T1 — Faber SMA-200 trend timing: the transplant fails, and fails instructively

    Cell Trades PF Return Strip-best Sum DD (acct) MC p95 Exposure B&H (DD) Verdict
    BTC W1 14 5.60 +315.0% −2.7% 45.8 (13.7) 64.3 50% +206.3% (71.9) ❌ strip-best, DD
    BTC W2 16 6.81 +146.2% +46.2% 10.5 (4.2) 19.3 59% +170.3% (53.0) ✅
    ETH W1 11 14.30 +887.9% −2.0% 61.0 (53.3) 63.7 55% +299.2% (80.1) ❌ strip-best, DD
    ETH W2 12 6.07 +100.5% +30.3% 9.9 (5.2) 17.5 48% +14.3% (67.6) ✅
    BNB W1 18 24.74 +1579.3% +119.8% 46.4 (18.1) 52.7 55% +1973.7% (76.1) ❌ DD
    BNB W2 33 1.34 +34.5% −73.3% 50.5 (50.5) 93.8 55% +107.3% (58.2) ❌ strip-best, DD

    The headline returns are enormous and they are worth nothing, which is exactly what the bar exists to detect. BTC's first window makes +315% across fourteen trades — and one trade, 29 Apr 2020 → 19 May 2021, is +317.8% of it. Delete that single hold and the window is negative. ETH's W1 is the same story at a larger scale: +887.9% total, +889.9% from one trade (22 Apr 2020 → 25 Jun 2021), strip-best −2.0%. Twelve of BTC's fourteen W1 trades are whipsaw losses of −2.6% to −10.1% — a 14.3% win rate, the second winner being an Apr–Sep 2019 hold worth +65.8%. The system is a lottery ticket on catching one bull run, wrapped in a year of small bleeding.

    The second windows pass cleanly on all counts, and that is the trap: run this on 2022–2026 alone and you would publish a 6.8 profit factor at a 10.5-point drawdown. The two-window gate is the only reason this isn't in the library.

    The comparison that matters, because the identical spec is entry #3 on SPY/QQQ: on US index ETFs it clears every gate in both windows with 9–15 point drawdowns. On BTC/ETH/BNB the same line produces 46–61 point drawdowns and a return one trade carries. The rule didn't change; the instrument did. A 200-day average on an index tracks a diversified thing that grinds; on a single crypto asset it tracks a thing that gaps 40% around the line and then whipsaws through it for months.

    T2 — one-week momentum: the effect is there, the drawdown is disqualifying

    Cell Trades PF Return Strip-best Sum DD (acct) MC p95 B&H (DD) Verdict
    BTC W1 53 1.98 +216.2% +156.4% 41.3 (38.2) 99.0 +269.3% (70.5) ❌ DD
    BTC W2 54 1.55 +82.8% +41.7% 30.7 (30.7) 77.2 +174.3% (51.8) ❌ DD
    ETH W1 48 2.38 +309.2% +216.3% 62.1 (62.1) 112.0 +427.9% (76.8) ❌ DD
    ETH W2 53 1.61 +112.2% +53.2% 40.9 (22.2) 87.7 +10.9% (67.1) ❌ DD
    BNB W1 51 3.95 +653.9% +279.8% 66.1 (45.3) 112.2 +2301.7% (72.1) ❌ DD
    BNB W2 53 1.56 +67.4% −9.4% 36.8 (29.0) 70.2 +82.3% (57.7) ❌ strip-best, DD

    This is the sweep's strongest signal and still a kill. All six cells are net positive with profit factors of 1.55–3.95 across ~50 trades per window, and five of the six survive strip-best (BNB W2 does not, at −9.4%) — the Liu–Tsyvinski effect, documented on 2011–2018 data, is still visible in two windows entirely out of sample for it. Then the drawdown column: 30.7 to 66.1 points against a ≤20 bar, with Monte-Carlo p95 tails of 70–112. Fee stress at 0.2%/side leaves every verdict unchanged (PF 1.45–3.82) — this is not a cost problem, and no cost assumption rescues it.

    The honest reading is that a weekly momentum filter is a real thing that roughly halves your exposure without halving your return, and it is not a ≤20%-drawdown strategy. ETH's second window is the one genuinely interesting cell — +112.2% against buy-and-hold's +10.9% — and it still carries a 40.9-point drawdown to get there.

    T3 — TTM squeeze: the lowest exposure in the sweep, still over the bar

    Cell Trades PF Return Strip-best Sum DD (acct) Exposure Verdict
    BTC W1 15 2.87 +113.2% +46.3% 24.4 (15.7) 21% ❌ DD
    BTC W2 17 1.77 +42.2% −1.8% 23.5 (15.5) 19% ❌ strip-best, DD
    ETH W1 18 3.72 +177.7% +106.1% 21.5 (16.2) 23% ❌ DD (near-miss)
    ETH W2 14 2.24 +57.9% +20.7% 17.8 (17.2) 14% ✅
    BNB W1 19 6.57 +460.5% +121.0% 45.2 (26.8) 23% ❌ DD
    BNB W2 20 1.61 +45.9% −29.5% 45.6 (24.5) 19% ❌ strip-best, DD

    The compression-then-expansion idea puts you in the market only 14–23% of the time and still breaches. ETH is the near-miss of the sweep — 21.5 in W1, 17.8 in W2, so one window over the line by 1.5 points — but that is a single asset failing a single gate, which is a near-miss, not an entry. BTC and BNB both fail strip-best in the second window on top of the drawdown. The 2026-08-01 "crypto vol-squeeze breakout" leftover is now closed at 1d, as its 1h cousin was in sweep #1.

    Informational — what the perp costs a daily trend trader

    T1 re-run on the perpetual (0.05%/side taker, funding charged over the holds) is the number a futures trader actually needs, and it is large: BTC W1 paid 50.4 points of funding, ETH W1 paid 64.7, BNB W1 paid 29.3 — against total returns of 266–1552 points, so 3–20% of gross, but on the losing whipsaw trades it is pure additional bleed. BNB's second window is the exception that proves the mechanism: funding there was net negative, paying the long 12.3 points. The lesson for anyone holding a multi-month crypto trend position: spot and perp are not interchangeable, and on a 385-bar hold the perp charges you roughly a fifth of a bull run.

    Verdicts

    # Hypothesis Verdict
    T1 Faber SMA-200 timing, crypto 1d Killed — 46–61 point drawdowns, and W1's entire return is one trade on both BTC and ETH. The identical spec passing on SPY/QQQ (entry #3) is the control
    T2 One-week momentum, crypto 1w Killed on drawdown — the published effect is visibly alive out of sample (PF 1.55–3.95, five of six cells strip-best-surviving) at 30.7–66.1 points of drawdown. Not a cost problem; fee stress changes nothing
    T3 TTM-squeeze breakout, crypto 1d Killed — 17.8–45.6 points at only 14–23% exposure; ETH is a 1.5-point near-miss, BTC/BNB also fail strip-best in W2. Closes the 2026-08-01 vol-squeeze leftover at the daily bar

    Neighbors were not run: the pre-registration allows them only for a family with genuine signal on ≥ 2 assets, and T1 had none (its returns depend on single trades). T2 reached the fee-stress phase and failed it identically.

    Structural conclusion — the second wall

    Sweeps #1–#7 measured the cost wall below the daily bar: sub-daily crypto edges are real and smaller than the fees. This sweep measures the drawdown wall above it: daily and weekly crypto momentum is real, survives strip-best, ignores fees entirely — and runs 30–66 points of drawdown, because that is what a single crypto asset does to anyone holding it directionally. Buy-and-hold's own drawdown in these windows is 52–80 points; trend systems cut that roughly in half, and half of 70 is still over the bar.

    Both walls have the same implication, and it is the one this library keeps arriving at: the crypto strategies that clear the bar do not hold directional exposure through the drawdown. They read positioning and get out (entry #1, entry #5), or they hold no direction at all (entry #6). Price-shape crypto is now tested and dead at 15m, 1h, 4h, 1d and 1w — the full ladder, both walls, published.

    A legitimate revisit of T1 or T2 needs a mechanism that addresses drawdown before seeing the results, not a filter added afterwards until the number looks acceptable — sweep #4 spent the Donchian family's revisit clause doing exactly that, and it still failed.


    Sweep 2026-09-18 (#10) — a second internals channel, and the ladder's top rung

    Pre-registered on 2026-09-18 and committed before the run (this section's git history is the timestamp; results are appended in a later commit).

    Question. Sweep #9 closed the price-shape ladder: crypto price shape is dead at 15m, 1h, 4h, 1d and 1w — of costs below the daily bar and of drawdown above it. Everything that has ever cleared the bar on crypto reads positioning: perp funding, directionally in entries #1 and #5, market-neutrally in entry #6. So: is there a second positioning channel that clears the bar — and does the funding mechanism itself survive at the one rung of the timeframe ladder never tested?

    The new data channel: the cross-exchange premium. Coinbase quotes BTC and ETH against US dollars to a mostly US, mostly spot, mostly institutional customer base; Binance quotes them against USDT to everyone else. The gap between the two prices — the "Coinbase premium" that desks watch — is a read on who is buying, in the same way funding is a read on who is positioned. It has never been tested here. It is free, it is deep (Coinbase daily candles reach back to 2016 for BTC and 2018 for ETH), and it is not derivable from the OHLCV the previous nine sweeps have exhausted.

    Disclosed confound, stated before the run: the two venues quote different assets — USD against USDT — so this series carries the USDT peg deviation as well as the venue premium. That is how practitioners compute it, and a depeg is itself a stress signal, but the number means "US spot demand and peg health", not a clean venue spread. No adjustment is made.

    Protocol. Engine: the app's own (runBuilderBacktest), harness committed at scripts/research/sweep10-crypto-internals.ts (public Binance + Coinbase data, cached, reproducible with no keys). Windows: W1 Aug 2018 → Aug 2022, W2 Aug 2022 → Aug 2026. M1/M3 run on BTC and ETH spot (Coinbase lists no BNB pair) at 0.1%/side, stressed at 0.2%; M2 runs on BTC/ETH/BNB perps at 0.05%/side futures taker with funding charged, stressed at the 0.1% spot-level tier. Gates per asset, unchanged: net positive in BOTH windows, strip-best positive in both, sum-point max drawdown ≤ 20. Cross-asset: ≥ 2 majors, or the README's narrow-clearer rule with first-class guardrails.

    How the premium reaches the engine. The DSL has no cross-exchange atom, so the signal is injected through the engine's aux channel — the same channel perp funding uses — and read with the funding_rate atom, with includeFunding false so nothing is charged as a cashflow. What is injected is the signal minus its threshold, so one operator (crosses_above 0) expresses both a fixed-zero and a moving-percentile trigger. Fills, exits, fees and every metric are the real engine, unmodified. If a premium hypothesis passes, shipping it as an adoptable entry requires a real DSL atom and a live feed — a build decision, stated here so a pass is not mistaken for something users can adopt the same day.

    Hypotheses (specs fixed before the run)

    # Hypothesis Spec Assets Stated prior
    M1 Premium momentum — US spot demand returning marks the turn Daily. Premium = (Coinbase close / Binance close − 1) × 100, 7-day mean. Long when it crosses above 0. Exit: 2×ATR(14) trail. Long only BTC, ETH Low–moderate. A zero line is an arbitrary level on a series whose mean wanders with the peg; the mechanism is real but the trigger is crude
    M2 Funding-squeeze at 1d — entry #1's frozen spec on the ladder's untested top rung Daily. funding_rate < 0 AND RSI(14) crosses above 30; 2×ATR(14) trail. Verbatim entry #1, only the bar changes BTC, ETH, BNB Moderate. The spec is cross-asset at 4h and BTC-only at 1h — the edge per trade grows as the bar does, so 1d should be the strongest rung. The risk is the opposite one: ~4–8 signals a year may be too few to distinguish from luck
    M3 Premium capitulation — the same "fade the crowded side, enter on the turn" shape as entry #1, on the new channel Daily. 7-day mean premium crosses back above its trailing 1-year 10th percentile — US selling was unusually extreme and is relenting. Exit: 2×ATR(14) trail. Long only BTC, ETH Moderate–high. This is the shape of the library's only durable directional crypto edge, applied to the only untested positioning channel. A trailing percentile also self-adapts: the premium's scale in 2018 and in 2026 are not the same number

    Both premium parameters are the natural round choices and are fixed here: smoothing 7 days ("one week"), percentile window 365 days ("one year"), percentile 10th ("the bottom decile"). The exit is the library's house exit for internals entries — the 2×ATR(14) trail entries #1 and #5 use — adopted unchanged rather than invented for this sweep.

    Neighbors run only if M3's primary shows genuine signal (positive and strip-best-surviving in both windows on ≥ 2 assets, drawdown ignored), and only as the pre-declared 5th and 20th percentiles — a plateau check, never a winner-picker. A monotone or flat neighbor grid supports the result; a knife-edge one kills it regardless of what the primary said.

    Decision rule, written down first. A full-bar passer on both majors ships as entry #7. A single-asset passer ships narrow, with its guardrails first-class. A premium passer ships as a research finding first and becomes adoptable only once the atom and the live feed exist. Anything else is a kill, published here.

    Results — all three killed, and the funding mechanism has a ceiling as well as a floor

    The premium data is clean: 3,302 of 3,302 bars align with Coinbase on both BTC and ETH, the 7-day mean premium sits at +0.014% (BTC) and +0.022% (ETH) and ranges −3.5% to +5.6%. The channel is real and well-measured. It does not clear the bar.

    M1 — premium momentum: right shape, wrong instrument class

    Cell Trades PF Return Strip-best Sum DD (acct) Exposure Verdict
    BTC W1 21 0.62 −31.8% −49.4% 34.7 (34.7) 11% ❌ return, strip-best, DD
    BTC W2 26 1.80 +51.9% +20.2% 29.3 (29.3) 24% ❌ DD
    ETH W1 26 2.42 +140.7% +58.6% 23.6 (23.6) 20% ❌ DD
    ETH W2 30 1.85 +69.9% +31.2% 21.0 (19.0) 26% ❌ DD

    ETH is the near-miss: both windows positive, both strip-best-surviving, profit factors 1.85 and 2.42 — and drawdowns of 23.6 and 21.0 against the ≤20 bar, so it misses by 1.0 and 3.6 points. BTC's first window is outright negative. A zero line on a series whose mean wanders with the USDT peg was the stated weakness of this trigger, and it behaved as the low prior expected.

    M2 — the funding-squeeze at 1d: the ladder has a ceiling

    Cell Trades PF Return Strip-best Sum DD Exposure Verdict
    BTC W1 6 1.55 +9.1% −3.5% 10.7 5.9% ❌ strip-best
    BTC W2 7 0.89 −3.0% −20.4% 11.5 3.4% ❌ return, strip-best
    ETH W1 3 2.57 +14.9% +2.3% 9.5 3.8% ✅
    ETH W2 4 0.00 −19.6% −19.6% 19.6 3.2% ❌ return, strip-best
    BNB W1 7 0.10 −47.9% −50.9% 48.2 6.1% ❌ return, strip-best, DD
    BNB W2 2 — +24.2% +2.1% 0.0 2.5% ✅ (2 trades)

    Dead, and for the opposite reason to 15m. At the daily bar the setup fires two to seven times per four-year window — BTC's entire W2 is seven trades, ETH's W1 is three. Nothing can pass a strip-best test on three trades, and nothing should: a single trade is 33% of the sample. BNB's first window loses 47.9% across seven trades. The drawdowns are fine (0–19.6 outside BNB) because exposure is 2.5–6%; there is simply no strategy there.

    That completes the ladder in both directions. One frozen spec, four rungs:

    Rung Verdict
    15m Dead — costs (sweep #7)
    1h BTC only, at ≤0.05%/side futures fees (entry #5)
    4h Cross-asset, the strongest form (entry #1)
    1d Dead — sample. Two to seven signals per four-year window

    The funding-squeeze is not "better on higher timeframes." It has a sweet spot at 4h, and it is bounded on both sides: by cost below, by sample above. Entry #1's timeframe is now measured from both directions rather than assumed.

    M3 — premium capitulation: the strongest prior in the sweep, and it failed

    Cell Trades PF Return Strip-best Sum DD (acct) Exposure Verdict
    BTC W1 13 0.80 −11.4% −28.0% 22.1 (20.0) 6.8% ❌ return, strip-best, DD
    BTC W2 18 3.32 +86.8% +39.8% 18.5 (9.6) 18.3% ✅
    ETH W1 10 3.91 +121.6% +42.1% 22.3 (19.5) 11.5% ❌ DD
    ETH W2 16 1.49 +31.8% −13.4% 33.4 (20.7) 12.5% ❌ strip-best, DD

    This was pre-registered with the sweep's highest prior — the exact "fade the crowded side, enter on the turn" shape that makes entry #1 work, pointed at the new channel — and it is the sweep's clearest kill. BTC's two windows disagree completely (−11.4% then +86.8%); ETH's do too, in the other direction (+121.6% then +31.8% with strip-best negative). Every one of the four cells that looks good has a partner that doesn't. Neighbors were not run: the pre-registration required genuine signal on ≥ 2 assets and there was none on even one.

    The stated prior was wrong, and writing it down beforehand is what makes that visible.

    Verdicts

    # Hypothesis Verdict
    M1 Cross-exchange premium momentum Killed — ETH misses the drawdown bar by 1.0 and 3.6 points with both windows otherwise clean; BTC W1 is negative
    M2 Funding-squeeze at 1d Killed on sample — 2–7 trades per four-year window. Establishes the mechanism's upper bound; 4h is a measured sweet spot, not an assumption
    M3 Cross-exchange premium capitulation Killed — window-disagreement on both assets, in opposite directions. The sweep's highest stated prior

    Structural conclusion — the channel is real, the drawdown is still the wall

    The cross-exchange premium is a genuine positioning signal: it aligns on every bar, it moves with stress, and pointed at ETH it produced two independent windows of profit-factor 1.85–2.42 that survive strip-best. It still fails, and it fails at the same place everything else in crypto fails — a drawdown in the low-to-mid twenties instead of under twenty.

    Sweep #9 found the same thing on price shape. That is now the consistent finding across eleven hypotheses and two data channels: on crypto, holding directional exposure long enough to earn the edge means holding it through 20–35 points of drawdown, whichever signal put you there. Perp funding at 4h and 1h is the exception because it exits fast and sits out 95% of the time.

    The honest summary for anyone asking this library for another crypto strategy: there are three, they are entries #1, #5 and #6, and they are three because directional crypto exposure and a ≤20% drawdown are close to mutually exclusive. Two further channels were tested to find that out.

    A revisit of M1 or M3 needs the premium as a real DSL atom with a live feed, and a drawdown-aware mechanism stated in advance — not a tightened trail chosen after seeing these tables.


    Sweep 2026-09-18 (#11) — is "internals beat price shape" a fact about crypto, or about markets?

    Pre-registered on 2026-09-18 and committed before the run (this section's git history is the timestamp; results are appended in a later commit).

    Question. Ten sweeps have established a pattern on crypto: strategies that read positioning clear the bar (entries #1, #5, #6), and strategies that read price shape do not, at any timeframe from 15m to 1w. The library has never asked whether that is a fact about crypto or a fact about markets. Equities have a positioning channel too — the volatility term structure, spot fear against three-month fear — and it has never been tested here.

    Why it deserves the slot. Every crypto candidate now dies on drawdown, because a single crypto asset does not offer a ≤20%-drawdown path to a directional edge. Equity ETFs do: three of the library's six entries live there. So this sweep takes the crypto thesis to the instrument class where the drawdown bar is reachable, rather than taking another crypto mechanism to a bar that eleven hypotheses have now failed to clear.

    The data. VIX and VIX3M daily closes from CBOE's public index archive (free, no key). VIX3M begins 18 Sep 2009, which fixes the history available: W1 Jan 2010 → Jan 2018, W2 Jan 2018 → Aug 2026 — deliberately the same windows as entry #3, so the results are directly comparable to a trend system on the same instruments. Stated limitation up front: the 2000–2010 stress decade that entry #3 and #4 report is not available for this signal, so there is no equivalent stress row and this write-up will not pretend otherwise. Each window does contain a major regime break (2011 and 2015–16 in W1; 2018, 2020 and 2022 in W2).

    Instruments: SPY and QQQ as the primary majors, IWM and EFA for breadth. Costs 0.03%/side with a 0.1% stress. Gates unchanged: net positive in both windows, strip-best positive in both, sum-point max drawdown ≤ 20, then ≥ 2 instruments.

    Correlation disclosure, stated before the run. V2 and V3 buy panic and sell the recovery, which is the same family as entry #2 and entry #4. If either passes it ships with the same one-sleeve warning those two carry, and the write-up will say plainly that a third dip-buy is not diversification. V1 is a timing system and belongs to entry #3's family instead.

    Hypotheses (specs fixed before the run)

    # Hypothesis Spec Stated prior
    V1 Term-structure regime timing — own the index while the curve is calm Long when VIX/VIX3M crosses below 1.0; exit when it crosses back above. No stop, condition exit — entry #3's shape, driven by internals instead of a moving average Moderate–high. Volatility clusters, and the inversions are the crashes. The risk is whipsaw: the ratio crosses 1.0 far more often than a 200-day average does
    V2 Term-structure capitulation — take the normalisation, don't hold the regime Same entry as V1; exit on a 2×ATR(14) trail — the library's house exit for internals entries (#1, #5) Moderate. A trail should cut the whipsaw cost that threatens V1, at the price of exiting good regimes early
    V3 VIX level capitulation (the control) Long when VIX crosses back below 30; 2×ATR(14) trail. The conventional panic line, same exit as V2 Moderate. Its job is comparative: if V3 ≈ V2 the term structure adds nothing over the level, and the honest write-up is about the level

    Thresholds are the conventional lines and are fixed here: ratio 1.0 (flat curve), VIX level 30 (the panic line). Neighbors, run only for a family with genuine signal on ≥ 2 instruments: ratio 0.95 / 1.05 and VIX level 25 / 35 — a plateau check, never a winner-picker.

    Decision rule, written down first. A full-bar passer on ≥ 2 instruments including at least one primary major ships as entry #7, with the correlation disclosure above if it is V2 or V3. A passer on IWM/EFA only is reported as a curiosity, not shipped — the library does not ship an edge that avoids the two most liquid names it tests. If V3 matches V1/V2, the finding is that the term structure is not the ingredient. Anything else is a kill, published here.

    Results — a strategy passed all five gates, so we built a sixth gate and it died

    This is the most useful sweep in this log, and it did not produce an entry.

    V2 cleared every check the vetting bar has ever asked for: net positive in both windows on three of four instruments, strip-best-surviving in all of them, drawdowns of 8.1–14.5 against a ≤20 bar, robust at 3.3× the fee assumption, and better when filled at the next open. It was one commit away from shipping as entry #7. Then a test that is not on the bar — is this signal better than entering on random days? — put its second window at the 75th percentile of pure chance, and it is not in the library.

    The rest of this section is that story with the numbers, followed by what it cost the rest of the library.

    The setup, as measured

    The term structure is inverted (VIX above VIX3M) on 324 of 4,242 trading days — 7.6% — with a median ratio of 0.885. So the entry signal fires as a genuinely uncommon event returning to normal, which is what made it worth testing.

    V1 — holding the calm regime: killed on drawdown

    Cell Trades PF Return Strip-best Sum DD (acct) Exposure B&H (DD) Verdict
    SPY W1 48 2.45 +75.9% +62.2% 12.7 (11.8) 92.5% +135.5% (19.4) ✅
    SPY W2 55 2.36 +99.3% +50.4% 30.5 (20.3) 90.4% +177.9% (34.1) ❌ DD
    QQQ W1 48 3.20 +126.1% +105.8% 13.0 (10.8) 92.5% +235.6% (16.4) ✅
    QQQ W2 55 2.25 +132.5% +60.4% 46.6 (27.3) 90.4% +334.1% (35.6) ❌ DD
    IWM W1 48 2.22 +87.5% +66.2% 22.4 (13.3) 92.5% +138.3% (29.4) ❌ DD
    IWM W2 55 1.63 +66.5% +36.1% 36.8 (24.9) 90.4% +89.2% (42.3) ❌ DD
    EFA W1 48 1.12 +10.6% −7.5% 26.2 (24.0) 92.5% +23.9% (27.3) ❌ strip-best, DD
    EFA W2 55 1.27 +26.8% −10.6% 36.9 (34.5) 90.4% +49.1% (38.2) ❌ strip-best, DD

    Standing aside only while the curve is inverted leaves you invested 90%+ of the time, which means you keep essentially all of the index's drawdown — 30.5 on SPY and 46.6 on QQQ in the second window. The 7.6% of days the signal avoids are not where the damage is. Killed.

    V2 — the same entry with a trailing exit: passes the bar

    Cell Trades PF Return Strip-best Sum DD (acct) MC p95 Exposure Verdict
    SPY W1 38 2.96 +68.3% +56.6% 12.0 (10.6) 16.9 32.1% ✅
    SPY W2 41 1.71 +26.9% +14.1% 9.4 (7.6) 18.0 20.0% ✅
    QQQ W1 38 3.99 +91.6% +77.6% 8.1 (6.4) 13.6 29.8% ✅
    QQQ W2 39 1.87 +44.0% +22.6% 12.2 (11.8) 24.2 20.6% ✅
    IWM W1 38 3.10 +83.7% +71.3% 10.7 (8.6) 17.4 28.0% ✅
    IWM W2 41 1.44 +23.7% +10.1% 14.5 (14.1) 27.5 19.1% ✅
    EFA W1 39 1.30 +17.2% +6.6% 14.4 (13.7) 28.9 23.8% ✅
    EFA W2 41 1.14 +6.2% −3.6% 10.5 (10.4) 25.3 19.9% ❌ strip-best

    Three instruments including both primary majors, every gate, ~240 trades across the passing cells — a better sample than any crypto entry in this library. At 0.1%/side the same three still pass. The difference between V1 and V2 is only the exit, which already tells you where the value is: not in knowing when to be invested, but in getting out on a trail.

    Neighbors (run because the primary showed genuine signal): threshold 1.05 passes SPY/QQQ/IWM with profit factors of 3.84–9.71 on 13–21 trades; threshold 0.95 fails on drawdown nearly everywhere (21.9–29.0). That gradient is monotone and mechanical rather than knife-edged — at 0.95 the trigger fires on ordinary calm-market noise, the trade count triples to 57–75 per window, and the drawdowns follow. Selectivity again.

    V3 — the control earns its keep

    Cell Trades PF Return Strip-best Verdict
    SPY W1 9 0.16 −15.3% −18.2% ❌
    QQQ W1 8 0.19 −13.0% −16.1% ❌
    IWM W1 8 0.17 −14.9% −18.0% ❌
    EFA W1 9 0.21 −17.4% −19.7% ❌
    SPY / QQQ / IWM / EFA W2 21–22 2.56–2.79 +36.6 … +52.8% + ✅ all four

    Buying when VIX crosses back below 30 — the line every financial television segment quotes — is negative on all four instruments in 2010–2018 at profit factors of 0.16–0.21, then positive on all four in 2018–2026. A textbook window flip, and the reason the control was pre-registered: the term structure is not a dressed-up version of the VIX level. Comparing a level against a fixed number fails; comparing spot fear against three-month fear does not. That comparison is the only part of this sweep worth keeping.

    The sixth test — and the reason there is no entry #7

    Before shipping V2 the obvious question was asked: is this a signal, or is it a good exit applied to a market that goes up? The test is cheap. Keep the window, the instrument, the 2×ATR exit and the fees; throw the signal away and enter on N random days, where N is the number of trades the real strategy took; repeat 400 times; see where the real return lands.

    Cell Real return Null median Null p95 Real percentile Reading
    SPY W1 +68.3% +9.5% +33.0% 100th Beats the null
    SPY W2 +26.9% +14.4% +50.4% 75.3rd Indistinguishable
    QQQ W1 +91.6% +18.2% +46.9% 100th Beats the null
    QQQ W2 +44.0% +24.0% +65.9% 79.5th Indistinguishable
    IWM W1 +83.7% +11.7% +47.1% 100th Beats the null
    IWM W2 +23.7% +9.5% +48.5% 72.3rd Indistinguishable
    EFA W1 +17.2% −6.7% +18.7% 94th Marginal
    EFA W2 +6.2% +4.9% +31.5% 54.3rd Indistinguishable

    In the first window the signal is unambiguously real — 400 random draws never once beat it on any of the three instruments. In the second window the same signal earns less than the 95th percentile of random timing on all four. Its +26.9% on SPY sits inside a null distribution whose own 95th percentile is +50.4%.

    A strategy that cannot beat a coin flip in its most recent window is not a strategy, whatever five other checks say. V2 is killed.

    Two things this is not. It is not the dip-buy in disguise: of V2's ~79 entries per instrument, only 1.3–5.2% are also an entry-#2 RSI(2) signal day and only 15–25% fall within three days of one, so these are mostly distinct days. And it is not a look-ahead artefact: the edge survives a full one-day signal lag and next-open fills independently (the combination of both leaves QQQ alone, which is disclosed above and was the floor test).

    The finding to keep: the volatility term structure genuinely predicted equity returns in 2010–2018 and stopped doing so in 2018–2026. Every V2 cell in the modern window is inside the noise. That is a result about the market, published here rather than sold as an entry.

    Verdicts

    # Hypothesis Verdict
    V1 Term-structure regime timing Killed — 90%+ exposure keeps 30.5–46.6 points of drawdown; the 7.6% of days it sits out are not where the damage is
    V2 Term-structure capitulation Killed by the null test, after passing all five bar gates, fee stress and both fill checks. First window 100th percentile, second window 72nd–80th
    V3 VIX level < 30 (control) Killed — negative on all four instruments in W1 (PF 0.16–0.21), positive on all four in W2. The control worked: the term structure is not the level

    The bar gained a sixth check, and the existing entries were re-audited against it

    The five original gates ask whether a strategy made money robustly. None of them asks whether its signal did anything. V2 proved a strategy can satisfy all five and still be chance, so the null test is now a standing gate — and the honest thing to do with a new standard is apply it backwards. Every shipped entry was re-run through it on the window its own page displays, using the committed configs (scripts/research/null-test-audit.ts, 400 draws per cell):

    Entry Window Real trades (exposure) Null trades (exposure) Real return Null median Null p95 Percentile Verdict
    #1 BTC funding 4h 2022–24 10 (1.9%) 10 (2.2%) +15.3% −0.2% +21.5% 90th ⚠️ above, not decisive
    #1 2024–26 12 (4.1%) 12 (2.6%) +13.6% −4.0% +11.3% 96.3rd ✅ beats the null
    #5 BTC funding 1h 2022–26 66 (2.3%) 65 (1.9%) +12.9% −4.8% +12.5% 95.5th ✅ beats the null
    #2 RSI(2) dip-buy 2018–26 76 (12.0%) 71 (7.1%) +36.4% +10.3% +30.6% 98.3rd ✅ beats the null
    #3 SMA-200 timing 2010–18 21 (80.6%) 9 (58.5%) +64.4% +46.2% +71.0% 89th ⚠️ above, not decisive
    #3 2018–26 29 (77.0%) 12 (60.3%) +76.5% +59.9% +93.9% 77.3rd ❌ indistinguishable
    #4 Double-7s 2010–18 55 (31.5%) 46 (17.2%) +56.4% +26.8% +47.7% 98.5th ✅ beats the null
    #4 2018–26 50 (28.5%) 42 (15.7%) +45.7% +32.5% +65.3% 76.3rd ❌ indistinguishable

    What holds up. The two funding entries and the RSI(2) dip-buy pass. The crypto rows are the most striking in the table: random 4h and 1h entries with the same trailing exit have a median return of −0.2% to −4.8%, and the funding signal turns that into +12.9% to +15.3%. Entry #1's first window is 90th rather than 95th, which is a caveat on a ten-trade sample, not a failure. Entry #2 sits at the 98.3rd percentile.

    What does not, and the honest reading of each.

    • Entry #4's second window (76.3rd). Its first window is 98.5th, so the mechanism was real and the modern window cannot be distinguished from random timing at the same exposure. That is now disclosed on the entry.
    • Entry #3 in both windows (89th, 77.3rd). This one needs a fair hearing rather than a headline: a return-based null is the wrong test for entry #3, whose published claim was never excess return. Its own write-up already says holding beats it on raw return and that the pitch is the drawdown column — a third of buy-and-hold's. What this table adds is that its return is not distinguishable from being long most of the time with a trailing exit. Both statements are now on the entry.

    One methodological caveat, stated rather than buried: the null usually realises fewer trades and less exposure than the real strategy, because a random signal that fires while a position is open is ignored. Less exposure means a lower null return, which biases every percentile in this table upward — in the strategy's favour. Entry #3's 58.5% null exposure against its own 80.6% is the extreme case. So these are floors on the percentile, not ceilings, and the two ❌ rows would be worse under an exposure-matched null, not better.

    Nothing has been delisted. The entries keep their evidence, their windows and their guardrails, and they now carry this result too — which is the only version of "we publish the failures" that means anything once a failure lands on our own shelf.