Holdout SetEvery hypothesis registered before it was scored. Every verdict published.

How a verdict gets made

Method

This describes how an idea becomes a verdict. It is deliberately a description of process rather than of any particular signal — the constructions that failed are published in full on theLedger, and the one that survived is not published at all.

1. The substrate

Two archives, both built rather than bought. A tick archive capturing the full depth-of-book for a fixed universe every session, accumulating daily. And roughly a decade of daily bars reconstructed from the exchange's own raw published files.

The daily archive matters more than it sounds. Reconstructing from raw exchange files rather than a vendor feed means delisted, merged and collapsed companies are still present — the names that vanish from a survivor-only dataset and quietly inflate every backtest run on it. Every split and bonus is carried in a hand-verified factor table with negative controls: a genuine collapse must survive adjustment, because a healed crash is corrupted history.

2. Registration, before any scoring

A hypothesis is committed to the repository with a date, and it states: the metric, the horizon, the bar it must clear, the capital the costs are computed at, the consistency requirement, and the conditions under which the registration voids itself.

That last one does real work. A registration voids if the signal fires on too few events to be meaningful, or on so many that it is not selecting anything. Two hypotheses here died that way without ever producing a verdict — which is the correct outcome, and one you cannot reach if you write the rules afterwards.

3. Cost hurdles are computed, not assumed

Every bar is expressed as a multiple of the round-trip cost at a stated capital — brokerage, statutory charges, and the bid-ask spread actually observed in the archive rather than a nominal figure. The same cost model is used offline and in simulated execution, so a result cannot be an artifact of the two disagreeing.

This is the single most common way retail research fools itself. A signal with a real effect that is smaller than its own trading costs is not a small opportunity; it is not an opportunity. One effect measured here is positive on every single day tested and still unusable, because it lives roughly a fifth of the way to the cost floor.

4. One scoring run

The hypothesis is scored once against its registered bar. Out-of-sample days are banked without being looked at, and a rehearsal on them is declined even when it would be convenient — data you have peeked at is no longer out of sample, and it never becomes so again.

A verdict is FAIL, VOID, or a qualified pass. It is recorded with its date and it is not revisited. Where a threshold was set badly enough to be worth re-registering, that is permitted once, stated in advance, and the second verdict is final for the whole family.

5. Forward paper, then a gate

A strategy that clears the gate does not get capital. It gets a pre-registered forward-paper protocol, published before its first observation, and it has to survive that on data which did not exist when the thresholds were written. The current protocol and its verdict schedule are on the Protocol page.

What this method has produced

23 registered hypotheses, 22 of them dead, one qualified pass awaiting its first real verdict, and zero strategies built on a signal that had not cleared its bar. Whether that is a good outcome depends entirely on what you think the alternative would have produced.

The full list is on the Ledger, and the counts on this page are computed from it rather than written down a second time.