One idea finally passed. Then I spent a week trying to break it.
On 16 July I tested three ideas about picking stocks and one of them passed.
It was the first thing to get through in six weeks of testing. I had by that point failed twenty-two ideas in a row, so I want to be clear that this was a good day.
It was also the start of a lot more work. Failing an idea takes an afternoon. Passing one means you now have to find out whether you fooled yourself, and that took most of a week.
The three I tested
All three went into the repository on the same day with the same four pass marks: beat a benchmark that just buys everything, by a stated margin; do not fall more than a stated amount at the worst point; beat the benchmark in most rolling twelve-month windows; and never lose to it by more than a stated gap in a single calendar year.
Two of them were defensive ideas. Buy the calmest stocks. Buy the stocks that move least with the market. Both did exactly what the textbooks promise. In bad periods they fell about half as much as the benchmark did.
Both failed anyway, and the reason is worth sitting with. The benchmark was rising fast over this period, and being calm is not much use when everything is going up. They protected me from a problem I did not have, and charged me the returns for it. Safe, and beaten by simply buying everything.
The third one passed. It is a long-term reversal idea: broadly, buying things that have done badly for a long time. I am not publishing how it is built, because I am still testing it and handing over the recipe would be a different kind of mistake. It cleared all four pass marks at both account sizes.
Then I tried to break it
This is the part I think matters more than the pass.
I checked the result harder than I check the failures. A result that tells me what I was hoping for is exactly the one I am least able to judge, so it got a full review: was the code only ever looking at past data, did the dates line up, did the internal evidence agree with what the code claimed to do. The pass survived all of that.
But the review found something uncomfortable. My written description of the strategy did not match the code. For weeks I had been describing my own strategy incorrectly in my own notes. The code was what shipped in the dated commit, so the code is what counts, and nothing had been quietly adjusted after seeing results. The result stands. I corrected the description and left both versions in the record.
I had been running a test I could describe wrongly. That is a small thing and it bothers me more than the failures do.
I built a stricter set of pass marks and ran it again. The first set judged things partly on how often they won. I replaced that with proper risk-adjusted measures, wrote the new marks down, and re-measured. The idea cleared the harder set as well. Two candidates I tested against the new marks failed.
I tried to make it safer and made it worse. The obvious next move was to mix in one of those defensive ideas as protection. I wrote that down as its own test and ran it. The combination was decent and it still failed, because its risk-adjusted return came out below the original on its own. The original already fell less in bad periods than the thing meant to protect it. There was nothing for the hedge to fix, so all it could do was water down the result.
That one surprised me. Adding protection is supposed to be the safe choice.
Why I am still not calling this good news
Two things, and I would rather say them than have someone find them.
The years it actually traded were unusually kind to it. The idea needs several years of history before it can pick anything, so for the early part of my data it simply sat in cash. The years where it was really trading were mostly a strong recovery period, which is close to the best possible weather for buying things that have fallen. I cannot tell from this data how much of the result is the idea and how much is the weather.
It is the most exposed to my assumptions about companies that disappear. Buying beaten-down stocks means buying more of the ones that eventually get delisted. My backtest assumes you get back part of the last traded price when that happens. That assumption is a guess, and this idea leans on it harder than anything else I have tested.
Neither of those makes the result wrong. Both mean a backtest cannot settle it.
What happens now
The idea does not get money. It gets tested forward on data that does not exist yet.
Since the start of August it has been running on paper, making its decisions on real dates against real prices, with the pass marks and the stop conditions published in advance. The first scheduled result is around February 2027. If it fails, that goes on the Ledger in exactly the same place a pass would.
I would rather have a result I can trust in February than a good story now. The whole point of the twenty-two failures was to build something that can tell me no. It would be a strange time to stop listening to it.