We built a gate, then failed it twice in one day — on purpose
On 14 July I built a scoring harness for swing-trading signals, registered four hypotheses against it, and ran them. All four failed. Earlier the same day, a separate ten-day gate had failed the best intraday signal I had.
Two gates, one day, nothing promoted. This was the system working.
The problem pre-registration solves
Here is the failure mode, and it is not stupidity — it is something close to the opposite.
You have a candidate signal. You measure it. It comes back at 1.7× your cost hurdle, and your bar was 2×. Now your mind gets to work, and it is good at this: 2× was always conservative. The sample includes an unusual week. The horizon isn’t quite the one you’d trade. Look at the top-conviction bucket in isolation — that clears easily.
Every one of those thoughts might be correct. That is exactly the problem. You cannot tell the difference, from the inside, between a legitimate refinement and a rationalisation, because they feel identical. The only reliable defence is to have written the bar down before you saw the number.
So that is the rule here: a hypothesis is committed to the repository with a date, and the commit states the metric, the horizon, the bar, the capital costs are computed at, the consistency requirement, and the conditions under which the registration voids itself. Then it gets scored once.
A threshold chosen after seeing the result is not a threshold. It is a description of the result.
What a registration actually contains
The bars are on the Ledger, attached to each entry, but the shape is worth spelling out. A registration is not just “must beat costs.” It has legs, and a hypothesis must clear all of them:
A level bar. The effect must exceed the round-trip cost hurdle by a stated multiple, at a stated capital. Not “be statistically significant” — significance is free with enough samples and pays for nothing. The bar is economic.
A consistency requirement. The effect must show up in most periods, not concentrate in a few. A signal that made all its money in one regime is a bet on that regime, and I do not know how to tell in advance whether the regime is coming back.
A structure requirement. Higher-conviction instances should perform better than lower-conviction ones, monotonically. If they don’t, whatever the pooled average is measuring, it isn’t the thing the signal claims to measure.
A void condition. This is the one people skip and it does real work. A registration voids itself if the signal fires on too few events to mean anything, or on so many that it isn’t selecting anything. Two hypotheses here died that way — recorded as VOID, not FAIL, because a miscalibrated instrument is not a market verdict.
The two failures
The intraday gate ran first. The candidate was a flow-divergence signal that had looked genuinely good in a five-day preview: positive at every horizon, positive on four of the five days, and strongest exactly where its own conviction measure was highest. It sat at 1.5–1.9× the hurdle against a pre-committed 2× bar.
Five more clean days had been banked in the meantime and never scored. I declined a rehearsal on them, explicitly, because data you have looked at is not out-of-sample data and never becomes so again.
When the full ten days ran, the five unseen days came in around −9 bps at thirty minutes and inverted the whole result. The conviction signature flipped: the bucket that had been the strongest in preview became the worst. Positive medians against negative means — small frequent wins, and catastrophic losses precisely where the signal was most confident.
A metric with 55–61% preview win rates was noise wearing a costume. Five unseen days and a bar I couldn’t move caught it before a line of engine code existed.
The swing gate ran that evening. Four hypotheses, registered by dated commit with fixed thresholds, scored once. None cleared. One voided on the degeneracy rule — it fired on roughly 45% of in-universe days against a 30% cap, which is not a signal, it is a description of the market being open.
What it cost, and what it bought
The cost is real and I want to be honest about it: weeks of work, and nothing to show that could be called a result. No strategy. No engine. A set of harnesses and a longer list of things that don’t work.
What it bought is the ability to believe the next answer.
The alternative programme — the one where the bar moves — would by now have an engine built on the flow-divergence signal. It would have looked good for a while, because the preview data was genuinely favourable. It would have failed later, with capital committed and much less clarity about why.
There is a version of this that sounds like humility and isn’t. The point is not that I am careful. The point is that I do not trust myself to evaluate my own results, so the evaluation is made mechanical and moved to before the result exists. The doctrine is a machine for making my judgement matter less.
One honest caveat
Pre-registration does not make you right. It makes you legible — to yourself later, and to anyone reading.
It cannot save you from a bad hypothesis, a flawed instrument, or a bar set in the wrong place. Two verdicts on the Ledger were later retracted because the underlying data had a defect the registration could not have anticipated; both the original and the correction are still published, because a record you can edit is not a record.
What it does is remove one specific failure: the one where you decide what counts as success after you know the answer. That is a small guarantee. It is also, as far as I can tell, the only one available.