Holdout SetEvery hypothesis registered before it was scored. Every verdict published.

Registered before the first observation

Forward-paper protocol

One strategy cleared the gate. This is what it now has to survive, and the entire point is that you are reading it before the result exists.

The protocol was committed to the repository in July 2026, before the first observation was recorded. The strategy's construction is not published — it is being traded forward, and publishing the recipe would be a different kind of mistake. What is published is every condition under which it will be judged, which is the half that can be gamed after the fact and therefore the half worth fixing in advance.

Tier 1 — Operational fidelity

Decision timing, fill timing, cost accounting, and isolation from the rest of the system must match the validated backtest exactly. These are not tolerances. Any deviation is a defect to be fixed, not a variance to be explained — if the paper book and the backtest can disagree, the paper book is measuring something else.

Orders are planned on one evening and filled on the next, at the following session's open, which is what the backtest assumed. Getting this wrong by a single day is one of the most common and least visible ways a paper result detaches from the study it claims to validate.

Tier 2 — Survivability tripwires

Hard kill

Breaching the registered drawdown limits on paper equity meanspermanent suspension. Not a pause, not a re-tune, not a parameter review. The strategy is finished and a replacement requires a new registration from scratch.

Soft flag

Exceeding the drawdown or the worst-year gap actually measured in the backtest triggers a mandatory written review. It does not stop the programme, and it does not permit a change to the strategy — it only forces the observation to be recorded.

Tier 3 — Scheduled verdicts

Two sittings, both dated in advance.

  1. Month six · ~February 2027

    Requires a minimum count of clean rebalances, no hard kill, and returns falling inside the backtest's own envelope. Passing opens the live-capital question and nothing else. It is explicitly not a claim of edge — six observations cannot support one, and the protocol says so in its own text so that a future version of me cannot pretend otherwise.

  2. Month twelve · ~August 2027

    The first genuinely out-of-sample twelve-month window, measured against the drawdown and edge figures the backtest produced. Honestly, this is a sample size of one.

Pre-rejected as evidence

Benchmark comparisons shorter than twelve months are excluded as verdict inputs, in advance. Over a few months the comparison is dominated by whichever way the market happened to move, and it will be tempting to cite if it flatters and dismiss if it does not. Removing it from the evidence set before knowing its sign is cheaper than resisting it later.