Reading a backtest sceptically, including ours
Historical results are produced by someone who already knew how the period turned out. Here is the checklist that applies to every one you will ever be shown.
A backtest applies rules to past data and reports what would have happened. Every one you see was produced by someone who already knew how that period unfolded, which is the fact that everything else follows from.
The checklist
Look-ahead
Did any rule use information unavailable at the time? Indicators needing future bars, values that change after the fact, or references to a higher timeframe's completed bar all qualify. See repainting.
Costs
Were spread, commission, and slippage included, and at what level? Omitting them can turn a losing method into a winning one on paper, and the effect grows with trade count. See liquidity and spread.
Fills
Does the test assume every order filled at the price shown? At levels where a lot of orders sit, a fill at the touch is optimistic. In thin instruments it may be fantasy.
Parameter count
How many adjustable values does the method have, against how many trades? A rule with eight settings and forty trades has enough freedom to fit almost anything. See curve fitting.
The test window
Why does it start and end where it does? A window beginning after a crash or ending before one is a choice, and the reason for it is rarely stated.
Out-of-sample
Was any data held back, and was it tested once? Testing, adjusting, and retesting turns reserved data into development data. After three rounds there is none left.
Data artefacts
On futures, was the test run on a continuous contract, and was it back-adjusted? Historical prices in a back-adjusted series never traded, so any rule referencing absolute levels is reading prices that did not exist. See rollover.
Survivorship
For equities, does the universe include names that were delisted or acquired? Testing only on companies that still exist selects for having done well.
The distribution, not the total
A headline return says very little. What matters is the shape: the worst drawdown, the longest losing run, and whether the result depends on a handful of outliers. A method whose profit comes from three trades in five years is a different proposition from one that grinds, even at identical totals.
Applying this to us
We have not published performance figures, win rates, or backtested results for True Range Research, and nothing on this site should be read as implying any. If we ever do publish historical results, apply every item above to them.
The general point holds regardless. Any historical illustration — ours included — was assembled with knowledge of how the period turned out. That is unavoidable for anything shown on past data, and it is the reason what you observe on your own charts, on the instruments you actually trade, should outweigh anything on a vendor's page.