Radar PereneRadar Perene
← home

Radar Perene / Archive / science

Honest backtesting: split the sample before testing anything

◦ Index methodology v2.2 (working papers with DOI). See the methodology.

Science

There is a genre of chart that never disappoints: the equity curve of a backtest. It climbs with the serenity of someone who knows how the movie ends — because it does. Tested on a past that has already happened, tuned until it works there, the rule displays an impeccable record right up to the day it meets an actual future. The contrast between the simulated curve and live performance is so routine it deserves listing as a stylized fact of the industry. And yet the problem is not the tool: testing rules against the past is one of research's legitimate instruments. The problem is whoever answers the question before asking it.

Honest backtesting is the testing of a rule on historical data conducted as a research protocol: hypothesis declared before any computation, data that respect what each date knew, sample split before the first test — one part for development, another untouched for verification — and full accounting of every attempt, including the ones that failed.

The definition describes a protocol, not a platform. The tutorials that dominate the term teach how to operate the software; almost none teaches how not to fool yourself — and the second lesson is what separates research from a shop window.

The protocol, piece by piece

An honest backtest is recognized by five commitments, all made before the first result.

The hypothesis comes before the computation. What will be tested, on which series, over which period, with which criterion of success — all written down before anything runs. This trail devoted an entire article to pre-analysis registration; the pocket version fits in one line: whoever sets the ruler after the jump has not measured the jump.

The sample is split before any peeking. One part of the data serves to develop and calibrate the rule; the other — out of sample — stays locked until the end and is visited a single time, as a verdict. The order is what matters: splitting after exploring everything is splitting nothing, because the researcher's memory has already contaminated both pieces.

Each date knows only what it knew. Financial statements enter on their publication date, not the fiscal one; revised series enter in the vintage of the time. It is the discipline of point-in-time data, described in look-ahead bias — without it, the test consults the answer key.

The universe includes the dead. Testing on the assets that exist today inherits survivorship bias: the bankrupt and the delisted left the photograph, and with them left the ugly half of the story.

Costs and attempts enter the bill. Brokerage, taxes and execution friction, which erode small patterns; and the total number of variations tried, because twenty silent attempts buy a "discovery" by pure chance — the mechanism this trail dismantled in p-hacking.

Why almost no public backtest follows the protocol

The short answer is that the protocol is expensive and the shortcut is invisible. Each commitment on the list costs time, data or pride — and violating any of them leaves no mark on the final chart. Two identical curves can hide opposite processes: one was born of a declared hypothesis and survived the locked sample; the other is the twentieth variation of an idea that failed nineteen times. The reader who sees only the curve cannot tell them apart — and the seller who shows only the curve counts on that.

Hence the practical criterion of this whole trail: a backtest's credibility is not in its result, it is in its biography. A test with modest returns, costs deducted, a complete universe and failures reported informs more than a spectacular curve with no birth certificate.

What the house practices — and deposits in public

This house uses backtests as a research instrument, not as sales material — there is no robot, no course and no signal on the shelf, which removes the structural incentive to embellish curves. The protocol described above is the one practiced internally, and its verifiable part is public: the data package accompanying the house's research — series, dated events and the prior registry of hypotheses — is deposited in an open repository with an active DOI (10.5281/zenodo.21399426), and the house's working paper series includes a study whose published conclusion denies its own tested hypothesis. In numbers: six working papers are public on Zenodo with their own DOIs, and one of them carries a negative outcome in the body of the text. Parameters and code of each test belong to the workbench; what can be audited from outside — the existence of the registry, the date, the willingness to publish the failure — is within any reader's reach.

The detail that tends to surprise: the most valuable piece of that arrangement is the uncomfortable one. An archive that publishes the tests that failed is the only external proof that the tests that passed faced the same ruler.

Frequently asked questions

Does a well-made backtest guarantee future performance?

No — and that is not a footnote caveat, it is the nature of the instrument. An honest backtest estimates how a rule would have behaved in a past regime; regimes switch, costs change, patterns shrink once known. The protocol reduces self-deception; it manufactures no guarantee.

How many years of data are enough?

There is no magic number: what matters is the count of independent episodes of the phenomenon under test. A rule about crises counts crises, not trading days — and twenty years of data may contain three relevant episodes. Samples that look long in days can be very short in events.

Does visiting the verification sample more than once invalidate the test?

Each visit spends a little of the set's statistical virginity: adjusting the rule after seeing the out-of-sample result turns verification into development in disguise. The rigorous standard is one documented visit; the realistic standard is declaring how many there were.

Why do commercial backtests rarely show failures?

For the same reason lotteries advertise winners: the selection is the product. A reader who wants a single criterion can use this one — ask for the list of what was tested and failed. The answer, or its absence, is the datum.

Here the statistics trail closes: from spurious regression to the protocol that keeps the pretty curve from lying. The next trail opens the toolbox of the researcher who works alone — starting with the trade's most underrated instrument, the reference library: Zotero for whoever researches markets without a department

House readings: today's note, in the Daily · the precedents, in the Atlas.

Submitting a specific rule to this protocol, from hypothesis registration to the out-of-sample verdict, is workbench material — the kind the house takes on request.

Read also: Look-ahead bias: the error that lets a model \"predict the past\"

This is the Radar’s memory. Today’s reading — regime, 5 lenses and the day’s analogs — is live, free.

Subscribe to Perene Semanal — US$ 29/mo →

See today’s reading →