Holdout SetI write down the test before I run it, then publish what happened.

The full list

The Ledger

Every idea I have tested, the pass mark it had to beat, and what happened. I write the pass mark into a dated commit before I run the test. If I picked it afterwards I would just be describing the result and calling it a test.

Tested
23
Failed
19
Thrown out
3
Passed
1

For the ideas that failed I show the numbers. They are finished, so the numbers cost me nothing. For the one that passed I show only the result, because I am still testing it and I would rather not hand over how it works. The counts above are calculated from the list below, so the two cannot drift apart.

2 Jul17 Jul114 Jul715 Jul916 Jul5FailedThrown outPassed
Every result, on the day it landed. 23 of them across fifteen days, because scoring an idea takes an afternoon once the equipment exists. One passed. Thrown out means the test itself was invalid, not that the market said no.

Intraday

12 tested · 0 passed

vwap_ext

FAIL

Bet that a price stretched far from the day average comes back

Written down on
Not written down first. This one came before I started doing that.
Result on
2026-07-02 · DR-21
Pass mark
Cost hurdle at target capital (4.2 to 22.8 bps depending on capital)

The signal I started with. Inside the range where I would actually have traded it, the edge was about 30 times smaller than my cheapest trading cost, and past five minutes it pointed the wrong way. I also ran it on a made-up price series that reverts perfectly by design, and it still lost money there. That told me the anchor was the problem. Killing this one is why I started writing tests down in advance instead of trying one idea at a time.

H1 stretch_cont

FAIL

Bet that a stretched price keeps going instead of coming back

Written down on
2026-07-06
Result on
2026-07-14 · DR-24
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

The opposite of vwap_ext: instead of betting the price comes back, bet that it keeps going. The earlier study had hinted this direction was stronger. It was not. Nothing in the first batch came close to the pass mark.

H2 ofi_cont

FAIL

Bet that heavy one-sided order flow keeps pushing the price

Written down on
2026-07-06
Result on
2026-07-14 · DR-24
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

When there is far more buying than selling sitting in the order book, does the price keep moving that way? No. Below the pass mark. Along with H7 this settles the question for me: the level of order flow does not tell you where the price goes next.

H1×H2 stretch_ofi

FAIL

Stretched price and order flow both pointing the same way

Written down on
2026-07-06
Result on
2026-07-14 · DR-24
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

Price stretched a long way, with order flow pushing the same way. Best of the first batch at +2.23 bps over thirty minutes, and it only worked on 4 days out of 10. Short of the pass mark on both size and consistency.

H3 ofi_diverge

FAIL

Bet on a reversal when order flow pushes against a stretched price

Written down on
2026-07-06
Result on
2026-07-14 · DR-24
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

Price stretched one way while order flow pushed the other way, so bet on a reversal. In a five-day preview this looked good: it worked at every time horizon, on 4 of the 5 days, and best of all exactly where the signal was strongest. It reached 1.5 to 1.9 times my trading cost against a pass mark of 2 times. I had five more days of data saved up that I had not looked at. I could have peeked. I did not. When I ran all ten days, those five unseen days came in around minus 9 bps and flipped the whole result. The cases where the signal was strongest turned out to be the worst ones. Many small wins, and a few very large losses right where I was most confident.

H4 orb_break

FAIL

Bet that a break out of the first-hour range keeps going

Written down on
2026-07-06
Result on
2026-07-14 · DR-24
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

Price breaks out of the first-hour range, so bet it keeps going. Below the pass mark, and I did not retry it. It taught me something separate: if you sample a condition every thirty seconds, one long stretch shows up as thousands of separate events. After this I stopped trusting raw event counts.

H5 rvol_mom

VOID

Bet that a trend continues when trading volume spikes

Written down on
2026-07-06
Result on
2026-07-07 · H5v2 precedent
Pass mark
≥2× round-trip cost hurdle · voids below 500 events

Trend continues when trading volume spikes. It found zero events, but not because the market said no. The volume input was a placeholder that had never been wired up and always returned the same number. I recorded it as thrown out rather than failed, and made that a standing rule: if my instrument is broken, that is not an answer from the market. A broken test gets one retry.

H5v2 rvol_mom

FAIL

The same idea, after I fixed the broken volume input

Written down on
2026-07-07
Result on
2026-07-14 · DR-24
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

The same idea, run again after I fixed the volume input. I kept the threshold exactly as it was, so I was testing the idea and not a new one. Below the pass mark.

H6 depth_imb

FAIL

Use the resting buy and sell orders in the book to predict direction

Written down on
2026-07-10
Result on
2026-07-15 · DR-30
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

Do the resting buy and sell orders sitting in the book predict the next few minutes? Yes. This is the most consistent thing I have found anywhere: it worked on all 10 days out of 10 at the short horizons, across 64,000 events, and the bigger the imbalance the better it did. Depending on account size it is somewhere between a twentieth and a third of what one round trip costs me. So it is a real effect that I cannot trade.

H7 ofi_accel

VOID

Use the change in order flow rather than its level

Written down on
2026-07-10
Result on
2026-07-15 · DR-30
Pass mark
≥2× round-trip cost hurdle · voids below 500 events

Not the level of order flow but the change in it. Only 12 events turned up in 10 days. My threshold was set so high that it sat right at the extreme tail of the real distribution. The test threw itself out under its own minimum-events rule instead of giving me an answer based on noise.

H7v2 ofi_accel

FAIL

The same idea, retried at a threshold measured from real data

Written down on
2026-07-15
Result on
2026-07-15 · DR-31
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

The one retry, spent on a threshold I calculated from the real distribution instead of guessing. That gave 1,860 events over 11 days, the most selective setting still allowed. The result was indistinguishable from noise, and the strongest cases were slightly negative. I am done with this family.

H8 roll_ext_diverge

FAIL

Same as H3, but measured against a rolling 30-minute average

Written down on
2026-07-10
Result on
2026-07-15 · DR-30
Pass mark
≥2× round-trip cost hurdle · ≥10 clean days · >1 regime family

The same idea as H3, but measuring the stretch against a rolling 30-minute average instead of the whole day. This is the closest I have come. Changing the reference point flipped it from negative to positive and it worked on 8 days out of 10 at every horizon. It reached 0.97 times my single trading cost, which is about half of what it needed to pass.

Swing

11 tested · 1 passed

S1 mom_12_1

VOID

Buy what rose over the past year, skipping the last month

Written down on
2026-07-14
Result on
2026-07-14 · DR-26
Pass mark
≥68.4 bps @20d · ≥60% of quarters · monotone quartiles · voids above 30% fire rate

Buy what has gone up over the past year, skipping the most recent month. It fired on about 45% of days against a 30% limit I had set. Something that triggers on half the market is not picking anything, so the test threw itself out before it could produce a result.

S1v2 mom_12_1

FAIL

The same, but only the strongest cases

Written down on
2026-07-15
Result on
2026-07-15 · DR-28
Pass mark
≥68.4 bps @20d · ≥60% of quarters · ≥6/8 years · monotone quartiles

The same idea, but only the strongest cases, using a threshold I measured from the data. This gave the biggest edge I have ever recorded: about 3 times the pass mark, over 66,574 events and ten years, getting steadily better as the signal got stronger. It worked in 52.8% of quarters against a pass mark of 60%. I had already written down why that matters. If something makes all its money in one kind of market, it is a bet on that market coming back, and I cannot predict that.

S2 near_52w_extreme

FAIL

How close a stock is to its 52-week high or low

Written down on
2026-07-14
Result on
2026-07-15 · DR-27
Pass mark
≥68.4 bps @20d · ≥60% of quarters · monotone quartiles

How close a stock is to its 52-week high or low. It beat the size pass mark but worked in only 18 quarters out of 36, which is a coin flip. Failed on consistency. An earlier run also showed the strongest cases doing worst, but that turned out to be bad data on my side and I took it back.

S3 weekly_pullback

FAIL

Buy a sharp one-week drop inside a longer uptrend

Written down on
2026-07-14
Result on
2026-07-15 · DR-27
Pass mark
≥68.4 bps @20d · ≥60% of quarters · monotone quartiles

Buy a sharp one-week drop inside a longer uptrend. Well below the pass mark. Worth saying what I got wrong here: my first run showed a very large loss, which I later traced to stock splits I had not adjusted for. An ordinary 1-for-10 split looks exactly like a 90% crash in raw data, so the test was buying fake dips that never recovered. On clean data it is simply flat.

S4 vol_breakout

FAIL

Buy when price movement suddenly gets bigger

Written down on
2026-07-14
Result on
2026-07-15 · DR-27
Pass mark
≥68.4 bps @20d · ≥60% of quarters · monotone quartiles

Buy when price movement suddenly gets bigger. It reached 0.99 times the pass mark, worked in 19 quarters out of 39, and the strongest cases were negative. Close, which does not count when you fixed the mark in advance.

Portfolio momentum

FAIL

Run the momentum signal as a real monthly portfolio

Written down on
2026-07-15
Result on
2026-07-15 · DR-29
Pass mark
Edge ≥+5 pp vs benchmark · maxDD ≤40% · ≥55% rolling-12m · worst-year gap ≥-15 pp

Taking the momentum signal and running it as a real monthly portfolio instead of measuring single events. Over 118 rebalances it failed all four pass marks at both account sizes, and it lost to a benchmark that just buys everything and costs nothing. This is the lesson I keep coming back to. A big average edge measured across thousands of overlapping windows does not survive being turned into an actual book you hold for a month at a time.

F1 low_vol

FAIL

Buy the calmest stocks

Written down on
2026-07-16
Result on
2026-07-16 · DR-32
Pass mark
Edge ≥+5 pp vs benchmark · maxDD ≤40% · ≥55% rolling-12m · worst-year gap ≥-15 pp

Buy the calmest stocks. It did what calm stocks do, falling about half as much as the benchmark in bad periods, and it badly lagged a benchmark that was rising fast. Safe, and beaten by simply buying everything.

F2 low_beta

FAIL

Buy the stocks that move least with the market

Written down on
2026-07-16
Result on
2026-07-16 · DR-32
Pass mark
Edge ≥+5 pp vs benchmark · maxDD ≤40% · ≥55% rolling-12m · worst-year gap ≥-15 pp

Buy the stocks that move least with the market. Same shape as F1 and the same result. I later tried both of these as a hedge alongside the one strategy that passed. That failed too.

S5 regime_momentum

FAIL

Momentum, but sit in cash when the market is below its long-term average

Written down on
2026-07-16
Result on
2026-07-16 · DR-34
Pass mark
Calmar ≥0.5 and ≥2× bench · edge ≥+5 pp · maxDD ≤40% · worst-rolling-12m ≥-25%

Momentum, but sit in cash when the market is below its long-term average. The filter did exactly what the textbooks promise on the one crash it was built for. Then it whipsawed in and out over two other years and gave back more than it had saved. It cut the losses and the returns together. That closes momentum for me at all three ways I have tried building it.

S7 ltr + low_vol

FAIL

Mix a defensive sleeve into the one strategy that passed

Written down on
2026-07-16
Result on
2026-07-16 · DR-34
Pass mark
Gate v2 legs, plus sleeve correlation ≤0.5 and combined Calmar ≥ the book alone

Mixing a defensive sleeve with the one strategy that passed. The combined version was actually quite good and still failed, which is the interesting part. Its risk-adjusted return came out below the original on its own. The original already had a smaller drawdown than the thing meant to protect it, so the sleeve had nothing to fix and could only water it down.

F3 ltr

QUALIFIED PASS

Long-term reversal. How it works is not published, since it is still being tested.

Written down on
2026-07-16
Result on
2026-07-16 · DR-32
Pass mark
Edge ≥+5 pp vs benchmark · maxDD ≤40% · ≥55% rolling-12m · worst-year gap ≥-15 pp

Passed all four pass marks at both account sizes, and later cleared a second, stricter set that I designed afterwards. I say passed with caveats for two reasons. Most of the years it was actually trading were an unusually good period for this kind of strategy. And of everything I have built, this one leans hardest on my assumptions about companies that get delisted. It is now being tested forward on new data, against a protocol I published before the first day of that test. First result due around February 2027.