Holdout SetI write down the test before I run it, then publish what happened.

How I built survivorship-free NSE data from free bhavcopy files

If you backtest Indian stocks using data pulled from a convenient API, your results are probably wrong, and wrong in the direction that makes you happy. Not by a little. By a lot, and in a way you cannot see.

This is about why, and about rebuilding the data so it does not happen. The source turned out to be free and public, which was slightly annoying.

Two problems you cannot see

Start with the obvious approach. Take the stocks in the index today, pull their price history, run your backtest.

Two things have already gone wrong.

The dead companies are missing. Your list only has companies that still exist and still trade. Firms that went bankrupt, got delisted or collapsed are just absent, not flagged as missing, as though they never listed at all. Any strategy that would have bought those names gets measured as though it never could have. Strategies that buy beaten-down stocks suffer worst, because the beaten-down stocks are exactly the ones that disappear.

You are using tomorrow’s list. Applying today’s index membership to old data means your 2018 backtest is trading the companies that turned out to be successful enough to still be in the index in 2026. You have handed your strategy next week’s newspaper.

Neither of these shows up as an error. No nulls, no warnings, no gaps. The backtest runs cleanly and gives you a number that is too high, and you have no way of knowing by how much.

I started on a dataset with both problems: about a million rows, ten years, and the current members of a broad index, pulled from a free API. It was fine for building the machinery and useless for judging whether anything worked, and I labelled everything that came out of it accordingly.

The fix is the exchange’s own daily file

NSE publishes a bhavcopy every trading day. It is a plain file listing every security that traded that session, with open, high, low, close and volume. They have been doing this for decades and the archive is public.

The idea is almost too simple. Build your history forwards from the daily files, instead of backwards from today’s index, and both problems disappear at once. A company that traded in 2017 and collapsed in 2019 is in the 2017 files, because it traded in 2017. Nothing has to be reconstructed, because nothing was ever taken out.

My first rebuild came to about 2.5 million rows across roughly 2,550 stocks and seven years, at no cost. Later I filled a gap in it and extended it forward, which took it past four million rows and a full decade.

How to check it worked: look for the dead companies. The test is “are the bankrupt ones in here”, not “does the file load”. I wrote tests that check for specific collapsed companies by name: DHFL, Jet Airways, Reliance Capital, RCom, RPower, Yes Bank, Sintex. Those tests failed on purpose against the old dataset and only went green on the rebuilt one.

Write the test that fails on your current data first. Otherwise you have written a test that says everything is fine.

The part that will actually cost you a month

Raw bhavcopy is not adjusted for corporate actions. This is where the free data stops being free.

When a stock splits one-for-ten, the raw close drops from 2,000 rupees to 200 overnight. Nothing has happened. Every shareholder is exactly as rich as before. But your data now contains a 90 percent one-day fall, and anything you calculate across that date is wrong: momentum, volatility, drawdown, all of it.

A scan of my raw archive found 220 of these cliffs. Each one needs an adjustment, and each adjustment needs checking, because getting one wrong is worse than doing nothing. A bad adjustment corrupts the data quietly, in a direction you cannot predict, and everything downstream inherits it.

The tests matter more here than anywhere else, and the important ones work backwards. It is easy to write a test saying “after adjustment, the HDFC Bank split is smooth”. It matters far more to write one saying Yes Bank’s collapse must survive adjustment. That was a real 80 percent fall. Real money was really lost. If your cleaning code smooths that out, you have not cleaned your data, you have deleted history. And you will never notice, because data that looks clean is exactly what you were hoping for.

I ended up with 179 checked adjustments and 79 places where I treat the series as broken rather than adjusting it. Demergers go in that second group, because the company before and the company after are two different businesses, and pretending they are one is just making things up.

The tests found a hole in my own data

Here is the part I did not expect.

The detector flagged three stocks with a cliff on the same day, at ratios that are not normal split ratios. Three unrelated corporate actions on one day does not happen. So either the detector was broken or my data was.

My data was. There was an eight and a half month hole in 2022 where the original download had quietly failed. The two dates either side of the gap sat next to each other in the file, so anything calculating a “one day return” across that point was actually spanning eight months.

The evidence had been sitting in plain sight the whole time. The dataset recorded its own row count and its own date range, and dividing one by the other implies about 170 missing sessions. Nobody had divided.

Two things changed after that. Gaps became proper data in their own right, recorded in a file, with a rule that no calculation may span one. And the diagnostics stopped assuming that two rows next to each other are two days next to each other.

The lesson goes beyond this project. Tests written against real data find the problem you never thought to look for. I wrote that suite to catch bad split adjustments. It caught my dataset instead.

What this costs you

About a month if you are careful. Most of it is corporate actions, and most of that is checking rather than coding.

What you get is ten years of history with the dead companies still in it, stock lists that do not know the future, and the thing that turned out to matter most: a reason to believe your own backtests. Nineteen of the twenty-three ideas on my Ledger failed, and I can be reasonably sure they failed for real reasons rather than because my data was lying to me.

Two things I would say plainly. The exchange’s raw files beat the convenient API not because they are better data, but because they are unprocessed data, and the processing is where the bias got in. And none of this is clever. It is the boring version done properly, which is why almost nobody does it, and why flattering backtest numbers are so common.

Traces to DR-13, DR-25, DR-12, DR-27 in the working repository, written at the time rather than reconstructed for this post.