Holdout SetI write down the test before I run it, then publish what happened.

I set the pass mark first, then failed it twice in one day

On 14 July I built the scoring code for swing trading ideas, wrote down four ideas to test against it, and ran them. All four failed. Earlier that same day, a separate ten-day test had failed the best intraday signal I had.

Two tests, one day, nothing passed. That was the system working.

The problem this solves

Here is how it goes wrong. Stupidity has nothing to do with it. If anything it is the opposite.

You have an idea. You measure it. It comes back at 1.7 times your trading costs, and you had said 2 times. Now your brain gets to work, and it is good at this. Two times was always a bit conservative. That sample includes an odd week. The time horizon is not quite the one you would trade anyway. Look at just the strongest cases, those clear it easily.

Any one of those thoughts might be correct. That is the whole problem. From the inside you cannot tell a sensible adjustment apart from an excuse, because they feel exactly the same. The only defence I know of is to have written the number down before you saw the score.

So that is the rule here. Before I run anything, the idea goes into the repository with a date. It says what I am measuring, over what horizon, what pass mark it has to beat, what account size the costs are based on, how consistent it has to be, and what would make the test invalid. Then it gets run once.

What I write down

The pass marks are on the Ledger, attached to each entry, but the shape is worth spelling out. It is never just “must beat costs”. There are four parts and an idea has to clear all of them.

How big. The effect has to beat what a round trip costs me by a stated multiple, at a stated account size. Not “statistically significant”. You can buy significance with enough samples and it pays for nothing. The question is whether it covers the cost of trading it.

How often. It has to work across most periods, not pile up in a few. If something made all its money in one kind of market, it is a bet on that market coming back, and I have no way to know whether it will.

Does it get better when the signal is stronger. The strongest cases should do better than the weakest ones, in order. If they do not, then whatever the average is picking up, it has nothing to do with what the idea claims.

What would make the test invalid. People skip this one and it does real work. A test is invalid if the signal fires too rarely for the answer to mean anything, or so often that it picks out nothing at all. Two ideas here ended that way. I record those as thrown out rather than failed, because a broken instrument is not an answer from the market.

The two failures

The intraday test ran first. The idea was about order flow, and it had looked really promising in a five-day preview: positive at every horizon, positive on four of the five days, and strongest exactly where the signal was strongest. It sat at 1.5 to 1.9 times my costs against a mark of 2 times.

Five more clean days had been recorded in the meantime and never scored. I turned down the chance to peek at them, on purpose, because once you have looked at data it stops being fresh and it never becomes fresh again.

When I ran all ten days, those five unseen days came in around minus 9 basis points at thirty minutes and flipped the result. The cases where the signal was strongest went from best to worst. Positive middles with negative averages, which means lots of small wins and a few very large losses, right where I had been most sure.

The swing test ran that evening. Four ideas, each written down with a fixed pass mark, each scored once. None passed. One was thrown out because it fired on about 45 percent of days against a 30 percent limit. Something that triggers on half the market is telling you the market is open, not picking stocks.

What it cost and what it bought

The cost is real and I want to be honest about it. Weeks of work and nothing at the end that anyone would call a result. No strategy. No engine. Some equipment and a longer list of things that do not work.

What it bought is being able to believe the next answer.

Run this the other way, where the pass mark moves, and by now I would have a strategy built on that order flow idea. It would have looked fine for a while, because the preview data really was good. It would have fallen apart later, with money in it, and I would have understood much less about why.

There is a version of this that sounds humble and is not. I am not saying I am careful. I am saying I do not trust myself to judge my own results, so I made the judging mechanical and moved it to before the result exists. The rules are there to make my opinion matter less.

One thing this does not do

Writing it down first does not make you right. It makes you easier to check, by yourself later and by anyone else reading.

It cannot save you from a bad idea, a broken instrument, or a pass mark set in the wrong place. Two results on the Ledger were later withdrawn because the underlying data had a fault the test plan could not have known about. Both the original and the correction are still up, because a record you can quietly edit is not much of a record.

What it does is remove one specific failure: deciding what counts as success after you already know the answer. That is a small guarantee. It is also the only one I have found.

Traces to DR-26, DR-24, DR-27 in the working repository, written at the time rather than reconstructed for this post.