My best backtest result was a bug, and fixing it reversed the answer
On the first day of recording a live feed, my bot logged about 106,000 signals and filled seven trades.
That gap bothered me for a week. A hundred thousand chances and I took seven of them. It looked like a system tripping over its own safety rules while the opportunity went past.
I was wrong about nearly every part of that sentence, and finding out took a bug, an outside reader, and a result I had to withdraw.
First, where the signals were going
I could not answer the question by staring at the code, so I made the bot explain itself. Every time it looked at a stock and decided not to trade, it now wrote down why, in a bucket.
The answer came on the first day of it. Of roughly every 409 signals that reached the risk check, about 408 were turned away for the same reason: I already had a position open, and my limit was one at a time.
The shape was clear. A flood of stocks not stretched enough to be interesting, a thin stream that were, almost all of them turned away because the single slot was full, and a handful of actual trades.
So the 106,000 was the same few over-stretched stocks being re-examined on every tick while I held one position, not a hundred thousand opportunities. On a fast feed, counting signals is counting how often you looked, not how much there was to find.
That much was solid, and it still is.
Then I measured what the limit was costing me
The obvious next question: if one slot is the bottleneck, what happens with more?
I replayed the day with the limit set to one, three, five, ten and twenty. The result looked clean and sensible. Going from one slot to three improved things. After three it flattened out. Three, five, ten and twenty all gave the same answer.
That flattening is exactly what a real effect looks like. You get the benefit of spreading across a few positions, then you run out of good candidates and adding more changes nothing. It is the shape you hope for, because it tells you where the sensible setting is.
I wrote it up.
Someone read the code
A few days later an outside review went through the engine line by line and found something I had not.
The count of open positions being handed to the risk check was being set from whichever instrument had just filled, not from the positions I had open, so it was only ever zero or one. The limit could never bind above one.
Every single run with a limit of three or more had actually run with no limit at all. That is why they all agreed with each other. The flat line I had read as “the effect saturates” was three identical uncapped runs sitting on top of each other. I had read the signature of a bug as the signature of an effect.
I fixed it, the engine started passing the real count, and I re-ran the same day.
The corrected answer pointed the other way
With the limit actually working, raising it made things monotonically worse.
One slot made a small profit on that day. Three lost money. Five lost slightly more. Ten lost the most of all, with fees three times what the single-slot run paid, because it traded three times as often.
The control worked, by the way: the one-slot run reproduced my original number to the paisa, which is how I knew the difference was the fix and not something else drifting.
So the conclusion inverted. The single slot I thought was holding me back had been the only thing protecting me. The second, third and fourth most stretched stocks were losers, not slightly weaker versions of the first one. My best pick made money on one favourable day and everything ranked below it gave it away.
Those 270,900 rejections I had been mourning were rejections of bad trades.
What I took from it
A result that flatters you is the one you are least able to judge. I did not check that sweep as hard as I check a failure, because it agreed with a story I already believed: that my risk rules were too tight. Now anything that tells me what I hoped to hear gets more scrutiny, not less.
Look at the shape, not just the numbers. Three settings producing identical output is not a finding. Real effects are rarely that tidy, and “suspiciously clean” should be a thing you check rather than a thing you enjoy.
The retraction stays up. This post is the record of both: the result I first wrote up and the correction that reversed it. I did not go back and quietly edit the first one. A record you can rewrite is not much of a record, and the withdrawn version is the more useful of the two anyway.
The whole thing lived for 48 hours and it only died because somebody read the code. That is the argument for showing your work to people who are not you, and it is a large part of why this site exists.